Pith. sign in

Paper Citation Record · LEDGER

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

As of 23 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 14 inbound Pith citation observations for arXiv:2505.20256.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20256 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:02:55.925858Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:07:21.383338Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T17:27:15.771664Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2232d5e1-3447-49bc-a637-195515e6dddc · outbound

This paper cites Baichuan-omni-1.5 technical report.arXiv preprint arXiv:2501.15368, 2025.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Baichuan-omni-1.5 technical report.arXiv preprint arXiv:2501.15368, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:50.909522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:50.909522Z digest=sha256:5fd34476e066070dba81dd1fa490863ac03f199d986ef82e263da016c4fa7c24

Observation d14e2908-41ad-4ccf-8550-d93ac69167d0 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:51.005221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:51.005221Z digest=sha256:87355e5775b0593b7dbad564c656aaf86066f93a1052a6c18634be26c210c4d2

Observation de29e52b-f9c2-48b2-9b1b-4832bbcb07e3 · outbound

This paper cites Attention-based multimodal fusion for video description.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Attention-based multimodal fusion for video description

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:03:00.035632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:51.127713Z digest=sha256:932d03b150c7a1b499b075626bae28ae95f0c86bbdc7a99a8bf29333e23d5130

Observation db73eb7c-7f3e-4be0-910c-c0d9f3f905d4 · outbound

This paper cites Merlot reserve: Neural script knowledge through vision and language and sound.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Merlot reserve: Neural script knowledge through vision and language and sound

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:59.778055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:51.236879Z digest=sha256:f964615262444ade829ad8ef14bab934563fa776bf6febda9b8059ce78cc3f77

Observation 1a800756-ad6a-4554-a341-e925a6bc26bc · outbound

This paper cites GPT-4o System Card.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration GPT-4o System Card

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:51.365763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:51.365763Z digest=sha256:4b8e1ecfa902eba5ee00747fcc39c4ea26ce20c8db597fea14a67bee8524688d

Observation 5a787639-7038-44a8-9486-9a03bd435645 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:51.438063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:51.438063Z digest=sha256:e0b10226dbf10d6140cb95b4c9c491e1b0430bc8c90d0bc6646b664282e9f4cb

Observation be76c2b1-c408-4058-bf43-522ab05a4069 · outbound

This paper cites From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:51.602440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:51.602440Z digest=sha256:f3430bc492311ef286e9a5019260819e1d78c71dd6977a2606d0883bc56f048c

Observation 2dba246a-4fb9-4b09-9aae-66b5f621b3a7 · outbound

This paper cites Mavors: Multi-granularity video representation for multimodal large language model.arXiv preprint arXiv:2504.10068, 2025.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Mavors: Multi-granularity video representation for multimodal large language model.arXiv preprint arXiv:2504.10068, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:51.665984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:51.665984Z digest=sha256:54b4a8534d7d443257243dd065f9acbb163ae0227cdf2134308b356b2037ddbe

Observation ec614a37-190e-4dd2-9939-6843c541cbfd · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Lisa: Reasoning segmentation via large language model

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:59.626240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:51.756900Z digest=sha256:a07e17490a2acc22f734a42ad1019165d6b661ca0a765c7e6a1ac792460c0af5

Observation 8b47e923-4132-4fb2-b4e6-cd716d3c4cb8 · outbound

This paper cites Qwen2.5-VL Technical Report.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Qwen2.5-VL Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:51.849583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:51.849583Z digest=sha256:3620a02e4f75b75aeaca9effd5ac82d0f5542938a42994a37ad5d000c8c26cea

Observation da2c69d7-9b43-48ea-8f86-2b0e4b4ecd34 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:51.955495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:51.955495Z digest=sha256:9879366cbb6f8f460038d5e21f14d6dff624bfa5418e9d3755f181d216d4a26d

Observation 89728534-4ffa-413d-875c-af011ef21c29 · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.067992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.067992Z digest=sha256:4f31c68252c1b13c70babd213ce4616d877c9d294d1cd7ec9b865bcfcf5dc39d

Observation ab764237-0230-4fb1-8c10-aa1c07c08872 · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.167671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.167671Z digest=sha256:90c69fffbcbef931c4168e7b495c466d6eb2dbe208c05d1ae48b97cee40c284a

Observation 8e81f7b7-2473-478f-ba07-1d29a9a0dee1 · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Glamm: Pixel grounding large multimodal model

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:59.476621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:52.272410Z digest=sha256:426c1597735b4cddc2addab44804f8e7ce1849bfc0b4b0040c6e888733f4d14e

Observation b90be1db-5ac2-4685-bb08-193e845dc734 · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Florence-2: Advancing a unified representation for a variety of vision tasks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:59.274294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:52.368454Z digest=sha256:0fe1429c82e05c58aaaea0b4624dcc751cca1f12609f1eb5d8e2b288f1889b0f

Observation e806b0ba-fb70-439e-bafa-89354ea7f852 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.459678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.459678Z digest=sha256:df772e13b078fbe1a07ddf18851852ec518938c19d50a39c4a781320ab6cb6d6

Observation 1768ab9d-764b-4ddd-92f2-d163ff7de1c4 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.541895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.541895Z digest=sha256:54e563b13cf316830bab25d1782d11159eded817762ed4247d67ceeeaa365100

Observation 0c77d780-6a4c-4f63-8797-8eaa616d254f · outbound

This paper cites an unresolved cited work.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:02:59.154520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:52.652896Z digest=sha256:4d4bec523834e5ac529fc2ba571739434d65d10f9d66c691cfefd6c4856b7c6d

Observation 4127edd4-722e-45d1-831c-5edaeb2445eb · outbound

This paper cites OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.741949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.741949Z digest=sha256:89a016b192149350a117498de84c990cc24091b4098004db92d38f379d8d2671

Observation 4ea6870f-df15-4826-9250-2e6d29f1f584 · outbound

This paper cites Ref- avs: Refer and segment objects in audio-visual scenes.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Ref- avs: Refer and segment objects in audio-visual scenes

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:58.977483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:52.854703Z digest=sha256:9213136f43ae13337b483e06d22e67ba512d23c100d3323e1b1ea61a699db860

Observation 7dc7196d-1d7d-4a5b-9b90-e1fa8f024dc9 · outbound

This paper cites VISA: Reasoning Video Object Segmentation via Large Language Models.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.905262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.905262Z digest=sha256:fe4080b504e6f8af4a6bd530799baa07b0c6e8e8c525e29878864cfa7b293b1e

Observation 0836b924-e8b9-452e-96a7-5763e205d746 · outbound

This paper cites GPT-4 Technical Report.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration GPT-4 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.016804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.016804Z digest=sha256:005ada50fcfdf0026f55b463b241e6609121eaf1fef6cd724b91bbb1e3ab2583

Observation 56cf3a5c-718e-402e-a6fa-5b117e9945c1 · outbound

This paper cites Qwen Technical Report.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Qwen Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.127369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.127369Z digest=sha256:877f01e8ce10a039df8581bb25fc4d752cf969f4b6a5c2dd9f57b03de7460b38

Observation 37f867b5-ad9a-4d51-9c38-7b886fb81541 · outbound

This paper cites DeepSeek-V3 Technical Report.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration DeepSeek-V3 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.175517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.175517Z digest=sha256:03a2c679b2c2c36ae83e5ad0523c2651ec9e98f1caa09fe7fb11d9de02b7689d

Observation fb06e6fa-3c67-441d-a998-be2940b55e26 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.295895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.295895Z digest=sha256:f72ca88f0704c5212cac1d5a3e1adc3c693607418e222714148cb3be3994ef41

Observation 373e72bc-1f8a-40de-bfaa-f65fa07922f0 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.381768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.381768Z digest=sha256:241e820b6261115d63996eb17b645b28fae9c8965a2b113832f745d541507fca

Observation e126d7be-81b3-4c64-988b-4074a4487086 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:58.820469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:53.464889Z digest=sha256:f9a3fe9e3124a9c5f1126ce0d9e5286ca8f09ed69928d022682af4cf3f033166

Observation 811d0ee5-5e8d-4291-b3e4-a069be9de93d · outbound

This paper cites Visual instruction tuning, 2023.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Visual instruction tuning, 2023

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.519220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.519220Z digest=sha256:bea625000f47e2d276a5c05b0fe59a5c8e96441cb0e22fc1456c31222a69fd6c

Observation 8a7f098b-4c37-4905-9a09-dca496a1a4fd · outbound

This paper cites Deepseek-vl: Towards real-world vision-language understanding, 2024.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Deepseek-vl: Towards real-world vision-language understanding, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:58.724157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:53.626176Z digest=sha256:1a5fdde7e22505e56dcc9938eb10d2223a4b71260790f1f5c9cb541a24bc4e1c

Observation 0fb3a83b-ac11-422f-95b0-30d323ede06c · outbound

This paper cites Qwen2.5-Omni Technical Report.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Qwen2.5-Omni Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.705119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.705119Z digest=sha256:cf8eba93d1527479c26195f444382821c79041760bb4ed0033a3a1735e266590

Observation 1596d63f-ea15-4435-9835-bba8ac980638 · outbound

This paper cites Omnibench: Towards the future of universal omni-language models, 2024.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Omnibench: Towards the future of universal omni-language models, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:58.455156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:53.786386Z digest=sha256:48e8f361526b18f035273aae940ae3f3f9bdb56e408f2150ae49b85d152f5b8e

Observation 7acc720e-527d-4dd3-ae3e-2dbdd88aae6c · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.893277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.893277Z digest=sha256:ed3ff8e4136ddb297b1fd3afb35f03ca438446b97531462c7776c64e32351007

Observation dc2370a3-20a0-475c-ab8b-8c05fdd5def9 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:53.986036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:53.986036Z digest=sha256:87ea2301dedfcf85bab7153f92928b2acdfbd42147c13d4a5ddd45b295c4cc2a

Observation ed74b537-c464-4e8c-ae1c-f4a2f2fb4fba · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:54.059626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:54.059626Z digest=sha256:9b4ec70b1829c571811f7488493ddb1f17bdf5cef1f6542dbc84c3eaf71fbf9c

Observation 5f7eead2-aa70-42dd-8e9d-189cd6d4bd3a · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:54.164163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:54.164163Z digest=sha256:fb9281d4920fc6b272a45754bc500f830e3bd0f451b01c8eedec7ae206475b25

Observation 84479aa3-2f2e-43ac-82ad-7c4c856e9879 · outbound

This paper cites R1-omni: Explainable omni-multimodal emotion recognition with reinforcement learning, 2025.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration R1-omni: Explainable omni-multimodal emotion recognition with reinforcement learning, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:58.301784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:54.253030Z digest=sha256:ee017d88a1e6848352300ea8c3b056f577c41d83a00edd18ec6fffa5b1f0be2d

Observation 76ba6c11-fdad-4b64-8791-6417ff75765f · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration SAM 2: Segment Anything in Images and Videos

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:54.312129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:54.312129Z digest=sha256:46da99e3c2477cef752597f7c26ac608c1f59ec8e8d2360b9d655176abca5eea

Observation 091839ed-40c2-4cb8-931b-5ad43caaca4f · outbound

This paper cites End-to-end object detection with transformers.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration End-to-end object detection with transformers

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:58.165672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:54.373043Z digest=sha256:97deb61e69befef3422eea384a8cdd6e69734ecb9b97b07175846946e4016b30

Observation ee01d021-eb3e-42d2-a4e6-7f0ab97fa304 · outbound

This paper cites MeViS: A large-scale benchmark for video segmentation with motion expressions.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration MeViS: A large-scale benchmark for video segmentation with motion expressions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:58.046533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:54.490953Z digest=sha256:a4621f33ac39243c556ac3214f95566e64cc49634972c6e51613d0763c1fe8f0

Observation 9a828aec-1762-4492-81f7-72aa9620e47a · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Generation and comprehension of unambiguous object descriptions

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:57.852661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:54.574343Z digest=sha256:b57e8c8c58ae032f9a176221760d779103cf6d732cf75c2557d0df25cf5e437b

Observation 2760fb1d-f276-4eb6-8b9e-6aae6a500311 · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:54.636835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:54.636835Z digest=sha256:0f523adc7d6659bd95b25dc8f28dee344e304d344d41f850f768bfbda60dc633

Observation b1dfba77-6695-479a-82fd-bed20612b408 · outbound

This paper cites Avsbench: A pixel-level audio- visual segmentation benchmark.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Avsbench: A pixel-level audio- visual segmentation benchmark

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:57.620689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:54.754197Z digest=sha256:4ab9e79265344dc1b082cef800aafc097f14c143bf6ba681e290893253fcb0fe

Observation d024c242-252c-4501-ae1f-6bcb9173d017 · outbound

This paper cites Avsegformer: Audio-visual segmentation with transformer.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Avsegformer: Audio-visual segmentation with transformer

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:57.429546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:54.861182Z digest=sha256:9d1b68e2a6d067c5ee1e7d635f89ca565316457c70de9d0522d85e1dc16db217

Observation 2121f94a-ed7e-481f-ac14-118ba329eea8 · outbound

This paper cites Prompting segmentation with sound is generalizable audio-visual source localizer.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Prompting segmentation with sound is generalizable audio-visual source localizer

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:57.234971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:54.935916Z digest=sha256:dfac445e8827a78f84687389c97efdf6ffa7d794064b3a1e0d8daec51b2a95ec

Observation fff28671-19d9-4f53-95d3-f7cedab4b6f5 · outbound

This paper cites Language as queries for referring video object segmentation.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Language as queries for referring video object segmentation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:57.080900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:55.054466Z digest=sha256:e830356e4ed82415a9023839215805bc8a9229e6479ebc07fb23264edae4b4c2

Observation b616ca86-1092-41be-9e2b-4c47089ceeb6 · outbound

This paper cites Robust referring video object segmentation with cyclic structural consensus.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Robust referring video object segmentation with cyclic structural consensus

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:02:56.912927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:55.159388Z digest=sha256:4c2f02db422dae1d0f076f88e501c79ede98b45fbc299a19599384d8bf156f4b

Observation 28b03452-ecb1-438a-bf2a-f2a1a82b46a0 · outbound

This paper cites TrackGPT -- A generative pre-trained transformer for cross-domain entity trajectory forecasting.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration TrackGPT -- A generative pre-trained transformer for cross-domain entity trajectory forecasting

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.274394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.274394Z digest=sha256:481423d6faecba330c3ef0b779f13d4638e1e8733f1d9407de2abe5f6d90b1fc

Observation 05bc131d-b4c1-4755-91d1-174dceae61f9 · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.376208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.376208Z digest=sha256:ddaf7f1ccc4d73ec3c1b9750f225fcfa56329c6712b45f591b7261c81d4cf5ca

Observation ed5cd32a-4c1b-4cbd-b21c-9fe3843808f9 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration LLaVA-OneVision: Easy Visual Task Transfer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.437771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.437771Z digest=sha256:117946bcf183eb9c0d336080124b31c53078cc1c122a4f4ce431dd9d26ee60d5

Observation a3fcd63b-ceaa-46f9-bd7f-2fcd422481f8 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.535106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.535106Z digest=sha256:97c851127ad2d2355c8856d5d2228e38059b86c979a4cedc1f285e2f87625e38

Observation 41704680-dbff-4445-9e9d-9142addf6d12 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.605882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.605882Z digest=sha256:758f5318be1eeb9ced3d32c9256060f18e586f0e1e5d82b3b0905b3a3ef7f8b0

Observation 54bdf31a-b38e-46dd-8339-8ec5817828e8 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.662704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.662704Z digest=sha256:59b0f8cc413a217d837b51c6bad5d5884df5da0250753ac3c02b730d66e6754e

Observation 30c68b51-6c35-4cf8-b986-ded5388dec21 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.721329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.721329Z digest=sha256:a2ce09d634a99e959ff8c70fde9360c11843d11c43da669efa2d722577223ae7

Observation e8162af0-fae5-4271-94fb-4f446fa0a774 · outbound

This paper cites an unresolved cited work.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:02:56.762742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T14:02:55.829573Z digest=sha256:a31394d4bec9f58d1addcde86b4c3d531147d9bd9f474dab76c8c05effb1402c

Observation ec44756a-9734-4793-b05a-a27563163f71 · outbound

This paper cites AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:55.925858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:55.925858Z digest=sha256:7195d5486dea28ffc8ae8c15bde03dcf9f7399ec040305d2c2bac7cac92b706b

Pith citing papers

Observation 748bcc57-ef43-4c6c-874a-623cafbcc7f9 · inbound

Group Relative Policy Optimization for Speech Recognition cites this paper.

Group Relative Policy Optimization for Speech Recognition Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T12:07:21.383338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:07:21.383338Z digest=sha256:450cd0434f515fc2322216c241538c17073e633c66ad5f1d7790e2d83b78bb71

Observation 11fa81bf-7cdf-4185-802d-a848fa6ccfb4 · inbound

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models cites this paper.

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:45:56.128382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T05:45:07.700571Z digest=sha256:509370c934a7df89344aedada29914dacef51e3c2167d2dac1835ad7a64d421d

Observation 9a349702-83c9-405d-ae4c-4abb8a12968e · inbound

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling cites this paper.

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:31:32.120343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T12:26:35.347190Z digest=sha256:d7ed276464e75c5c5ec8e570d707d1b1e023b97ad5b5030ff55e167d25c936d3

Observation 60a8a706-ac62-4312-9138-576e90d3bfca · inbound

Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning cites this paper.

Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:37:59.982181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T14:37:05.402850Z digest=sha256:325e3781ea153922c425a7b1084bfe5288b79ba33a2aaadf803599f22d333286

Observation a35bda6a-7128-40f8-b307-a518e00bcfc2 · inbound

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering cites this paper.

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:11:01.973351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T17:45:51.528645Z digest=sha256:7440b01bc8376cfb52e4054c5fb449a6ed3c6f07d2e2603c23b2316a129f1401

Observation 5cd62825-00f7-443b-b17b-2f24b32fc7a8 · inbound

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding cites this paper.

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:03.827672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T15:07:45.595260Z digest=sha256:38272cdc83e0c79995fa782ba6e19a0620f136ef72aa0d512dcfd6e4287a7579

Observation 3eb7c43c-ccbd-4e9a-a055-c0581c6ba1b4 · inbound

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models cites this paper.

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:10:28.298818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T14:10:03.707886Z digest=sha256:bf53345bed32548030a7cec22d66724629d7e1cdfcffdf2c14b5a0d57b2f101a

Observation 6eb1a9ba-ff0e-40fe-a774-7737a268bd6a · inbound

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models cites this paper.

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T21:18:46.566338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:18:46.566338Z digest=sha256:8fd2d751112befac9504be8bfbff86c570864bdae03d3e3300a731776ab4d9bd

Observation 33213efd-ece5-4758-826a-c2693ae4e7c5 · inbound

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs cites this paper.

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:22.082628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T12:05:54.551728Z digest=sha256:87946b8c1fc358c97b58ab9a460a82e57614fa8f2c4a3d79cc81d99a56a4152b

Observation 6baa4af0-3229-4b99-b6c1-dfea506983bf · inbound

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers cites this paper.

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:18:32.211166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T08:02:53.574120Z digest=sha256:00a5bc014060aba89a0bb33ade06053188e331679f71a469f64053db7e9c8a00

Observation cf8a1670-c68e-4524-ad51-1587028d1389 · inbound

PRIMED: Adaptive Modality Suppression for Referring Audio-Visual Segmentation via Biased Competition cites this paper.

PRIMED: Adaptive Modality Suppression for Referring Audio-Visual Segmentation via Biased Competition Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:30:53.325802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T02:30:44.486435Z digest=sha256:e4634820cb460c92f8777fdb97de614f9887a1cdf640f6d4e3f4fc1389268b78

Observation e46a2ae7-4b41-41ec-9c82-fcdb79ddfe19 · inbound

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation cites this paper.

RCoT-Seg: Reinforced Chain-of-Thought for Video Reasoning and Segmentation Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:30:59.904231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T01:16:25.031349Z digest=sha256:864f6d27889914e5bca9f49d80711eeffa0431c0a257cee5d33f81c0dd7f9710

Observation 59631231-82f1-48ef-91b3-68a6c3106ccb · inbound

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching cites this paper.

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.106597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T10:48:24.523702Z digest=sha256:260700fae62daed2c2ccbdab02983dd4e300d6ae36f7243998a2ea518f889bab

Observation 6c937d94-f391-4895-86bf-b614ce2e5c02 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

Reference 202

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.773107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:48d2219967174acd023dac74745b448c12f5ce6e8d25421ad1fd84dbc5c7dad5