Pith. sign in

Paper Citation Record · LEDGER

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs

As of 6 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 2 inbound Pith citation observations for arXiv:2601.13836.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.13836 v2

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T09:30:05.747416Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-11T01:53:01.939765Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T09:46:26.384867Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved71
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5342e9b2-2050-4423-834a-b4495beb432b · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T09:29:58.247681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:29:58.247681Z digest=sha256:dd4821e4ea9ebb1f39f2fc9986777c50641a5d0c48c469755434624fcfe04db6

Observation 117b96c1-1289-4c89-8fa5-b8541330e63b · outbound

This paper cites Claude haiku 4.5.https://www.anthropic.com/news/claude-haiku-4-5, 2025.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Claude haiku 4.5.https://www.anthropic.com/news/claude-haiku-4-5, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T09:29:58.342803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:29:58.342803Z digest=sha256:78e0d975eae2f94d609b2119b617ac504da59ce1f416abc40f631e37b510250e

Observation b9b5dc81-54ec-4d9e-8b4a-050b10c77247 · outbound

This paper cites Qwen3-VL Technical Report.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Qwen3-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T09:29:58.420522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:29:58.420522Z digest=sha256:ae47484964712715b0f934ca9669ac6af456078fb762c7078ae8d3ea9a333f19

Observation f93242b4-56c5-4950-9ebf-df4fbbcb86f4 · outbound

This paper cites Qwen2.5-VL Technical Report.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T09:29:58.557515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:29:58.557515Z digest=sha256:d08dec3909370112c43c9d8ecf906530ed91a0c92d79a5a66a47ee1e06755dd7

Observation 74decc1f-576f-4c53-bbbb-0510b7fcf13f · outbound

This paper cites JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T09:29:58.700322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:29:58.700322Z digest=sha256:c549c26897f5e394f4f02d57cf2dc11b116a12c1ef48662d060daf92b34ce170

Observation 291f06fe-e5cc-4f5a-a1a3-08be9955d49e · outbound

This paper cites Avocado: An audiovisual video captioner driven by temporal orchestration, 2025.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Avocado: An audiovisual video captioner driven by temporal orchestration, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T09:29:58.751135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:29:58.751135Z digest=sha256:a79d6b63d64a1bd37d226dc3b8ffd9d5eb87a975675175baa3d85bbb2c6ac3bd

Observation bdd9e760-8893-4b3a-9ab1-63ee9fc832ea · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T09:29:58.868545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:29:58.868545Z digest=sha256:bd701cddefdf3485076f4c888b72704600756604a7d07bafeab8ac3c54e0e3a5

Observation f94943c4-2e0a-4714-975b-c1967b14ffec · outbound

This paper cites AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T09:29:58.979284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:29:58.979284Z digest=sha256:e24e4f99190b007afa6d20a8ec75f806a5e16899c8203c11133b859c41ceeaa3

Observation 6b74ba86-77ec-4db6-8b20-4449c4c74f80 · outbound

This paper cites Qwen2-Audio Technical Report.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Qwen2-Audio Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T09:29:59.113453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:29:59.113453Z digest=sha256:363995a7141d195f7a16259476292b770dcf1da00239c84fea1ffd8ecc3f616f

Observation e4c8fac1-5cb4-430e-a2a6-bb14c36720b0 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T09:29:59.193365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:29:59.193365Z digest=sha256:773cffe093c24039f21335cf8a053b72eaa5a81d1e3e1e84f2925634fcc7293b

Observation 01b70bed-b4a6-4b00-888e-fbc0c43ce51b · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs An image is worth 16x16 words: Transformers for image recognition at scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T09:29:59.272939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:29:59.272939Z digest=sha256:57a6999fc3c8afdff6bd0ab469364bb645fe2b916e3f04023481b0d218ff8810

Observation 6dbf7de3-6058-46c9-952b-695a5ec7021a · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T09:29:59.347795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:29:59.347795Z digest=sha256:336929e1c6fca60875a127bc93d9356ce1c4b8d71629f900c9587dca7af6cbf7

Observation c5d976f3-30e0-4bad-a67f-a22ea71eed09 · outbound

This paper cites Longvale: Vision-audio- language-event benchmark towards time-aware omni-modal perception of long videos.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Longvale: Vision-audio- language-event benchmark towards time-aware omni-modal perception of long videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T09:29:59.418454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:29:59.418454Z digest=sha256:df6008bef364c735baf1926168d0ffcfe08581e638bc89f5b2892ec7e6e367aa

Observation ee42a785-38e9-4c48-a9ca-1219c9646491 · outbound

This paper cites Gemini 3 flash: frontier intelligence built for speed.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Gemini 3 flash: frontier intelligence built for speed

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T09:29:59.587619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:29:59.587619Z digest=sha256:1647986b4b4fafcfbff1ab630dcf48857cc466b0bbcb4f9367e29c6017ce4579

Observation 28da8c2f-0656-4530-b6f4-c4838e7d2a79 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T09:29:59.696384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:29:59.696384Z digest=sha256:ae605b5e654448e83f591db657b3c457596eae6300e1af96cd5ccffc98514056

Observation 10659347-8211-4457-b856-c13cc9598dfe · outbound

This paper cites WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T09:29:59.779406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:29:59.779406Z digest=sha256:6c6a9037f882182571459386d59ad7b98882ee7a23cdcf551926e15629872570

Observation 204b01f8-c23c-47b3-97fe-badefa6dd0d6 · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs LoRA: Low-rank adaptation of large language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T09:29:59.896596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:29:59.896596Z digest=sha256:6695210cac0ab0f7ff791cba8d2feb89dbab8bbf226d53753c20f1189f1ca9dd

Observation 815da43e-6d71-47a3-ad10-250e9834563a · outbound

This paper cites WavLLM: Towards robust and adaptive speech large language model.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs WavLLM: Towards robust and adaptive speech large language model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:00.164196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:00.164196Z digest=sha256:a30a263c40603e8a0fa3af9feade7454f74142893f415147acdf89bde02196c4

Observation 8d12d096-ae67-436b-95c2-c3d37645874d · outbound

This paper cites Forecastbench: A dynamic benchmark of AI forecasting capabilities.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Forecastbench: A dynamic benchmark of AI forecasting capabilities

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:00.271973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:00.271973Z digest=sha256:3b009423782f5f4a988c41a0738cca9c3ae94f62dd9c266a99cf738a969275ac

Observation 0145f447-cc1b-4b41-8253-502b0bac7606 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:00.354187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:00.354187Z digest=sha256:9b45248bb15396497fb49342edcc396d01848c5233830433c0069db7e0bba4cf

Observation 94a1e897-872f-4722-a3df-18de237e3985 · outbound

This paper cites What is more likely to happen next? video-and-language future event prediction.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs What is more likely to happen next? video-and-language future event prediction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:00.438253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:00.438253Z digest=sha256:86aa1fc785510f2d6d058e03e9915eff8bd23ec1b286fe8a92c153867a32b6e6

Observation e8dd5021-4dad-490f-9f84-a3c0e8855eb2 · outbound

This paper cites Omnivideobench: Towards audio-visual understanding evaluation for omni mllms, 2025.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Omnivideobench: Towards audio-visual understanding evaluation for omni mllms, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:00.515548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:00.515548Z digest=sha256:e0635b7eff6ed051521b3bcdddf4f4cef658dacc9decb65d8432f77a060b6552

Observation 34a15548-1bc1-49be-aa18-ddf1432f0031 · outbound

This paper cites Learning to answer questions in dynamic audio-visual scenarios.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Learning to answer questions in dynamic audio-visual scenarios

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:00.627809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:00.627809Z digest=sha256:ae95b2e1f46fa9e26d1538e511ee23688fcd94377390c7c5108f955a232b5488

Observation 6431bbed-44f9-4d53-8491-82ed7ec2121d · outbound

This paper cites an unresolved cited work.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:00.712306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:00.712306Z digest=sha256:33e706f9b236b784107097ff2708e24169d32d9c6f92588c24fb7a5870c6446f

Observation 563cf7eb-5382-4076-b190-ef4b60f00b4a · outbound

This paper cites Intentqa: Context-aware video intent reasoning.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Intentqa: Context-aware video intent reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:00.839550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:00.839550Z digest=sha256:548eae3317cdcc6d6656f69952bbf3cc9891d82958e5be7e268fe89899789cc6

Observation 182110fe-57b5-48cc-9470-3a8e95219cd5 · outbound

This paper cites BLIP-2: Bootstrapping language-image pre-training with frozenimageencodersandlargelanguagemodels.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs BLIP-2: Bootstrapping language-image pre-training with frozenimageencodersandlargelanguagemodels

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:00.927299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:00.927299Z digest=sha256:9108c2f99787a17ac33584accd7d2a2f62cda7f7945336b8b408f391fb229aa7

Observation a77102ca-79f2-4f92-b77d-b882fcbe62b3 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Llama-vid: An image is worth 2 tokens in large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:01.017385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:01.017385Z digest=sha256:3f3ce01deac800c9c9dfd7690c8871c9f0486d67953c8b9c5b4d5b90255c8c84

Observation a06cdbbc-e55d-41a5-a969-58a42594810e · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Video-llava: Learning united visual representation by alignment before projection

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:01.105450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:01.105450Z digest=sha256:85af5e605617824f62b1a03a38a02ccbf43bf07dfb83998431d823c43de9c878

Observation 8ed4ff97-65b9-4479-a7f4-d5dfaf9d1941 · outbound

This paper cites Visual instruction tuning.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Visual instruction tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:01.199391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:01.199391Z digest=sha256:91bacfa50ed881cf5e4c02f3dec98971fdcbd65a15da5ed8c2eedc31737e7f07

Observation 52c170a9-6577-4502-9730-ac4b1ecd25c6 · outbound

This paper cites Ola: Pushing the Frontiers of Omni-Modal Language Model.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:01.361938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:01.361938Z digest=sha256:3d462b10edd6c094ceb6f3aa762c144c78e091d299350a5d0efb041e081fe7c5

Observation 5fe89d08-cb9d-45d4-8a54-1ef956aff6a6 · outbound

This paper cites See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:01.498544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:01.498544Z digest=sha256:fa30d97d058c05d83c17395f6ee26dcd8b32bc3407f904eeb5d04fe2f59eeecf

Observation f4e04727-99ce-481b-8967-68008123f0ba · outbound

This paper cites GPT-4o System Card.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs GPT-4o System Card

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:01.633482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:01.633482Z digest=sha256:32ea104298204e39a9b8daf46e19630abe6a260c71eb40feb781e772b934f680

Observation 0d22d709-2fed-4727-b621-921d21e4624d · outbound

This paper cites Learning transferable visual models from natural language supervision.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Learning transferable visual models from natural language supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:01.732145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:01.732145Z digest=sha256:21f703e35e38cde20d4fbd506c6a68c54af5a3312b88a74b69639482e8845da7

Observation 456f5fc6-ad34-43ed-ad58-4a866a3fe9eb · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Robust speech recognition via large-scale weak supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:01.829896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:01.829896Z digest=sha256:58be1f65d774a1ccc12dba7466d841ce08a2e402eef50417207af1a0d34d5382

Observation be511dd8-3e82-4048-a917-08c72cf835c7 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:01.955904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:01.955904Z digest=sha256:9c10b8458dd1fb9da3bd824b7b485532abe74ab5a524df68de6ec35d4ee70b23

Observation db963ebe-3076-45e8-9ccc-49841ec3cbb7 · outbound

This paper cites SALMONN: Towards generic hearing abilities for large language models.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs SALMONN: Towards generic hearing abilities for large language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:02.029586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:02.029586Z digest=sha256:f3c9d6005d24c1843db5f233da2aa03b8d8518de9a50e4c420bf788f7188106a

Observation e2b5552f-7812-439e-a4ac-c51df590b3c9 · outbound

This paper cites video-salmonn 2: Caption-enhanced audio-visual large language models, 2025.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs video-salmonn 2: Caption-enhanced audio-visual large language models, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:02.081331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:02.081331Z digest=sha256:78dfc07a307ada15be95caf470989f003e5d6bc9229a6177933aac24db38fb4d

Observation 91eff93a-8fef-440f-88b3-520c1ba45c48 · outbound

This paper cites Empoweringllmswithpseudo- untrimmed videos for audio-visual temporal understanding.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Empoweringllmswithpseudo- untrimmed videos for audio-visual temporal understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:02.167133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:02.167133Z digest=sha256:ff6c2d9ad703bb8ef4b7b970da68dfc4390c56dbee8fd2ece0999bc96e0a6ed3

Observation 999cd9e9-7450-41b6-81d1-bffa35dbba45 · outbound

This paper cites Fostering Video Reasoning via Next-Event Prediction.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Fostering Video Reasoning via Next-Event Prediction

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:02.227226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:02.227226Z digest=sha256:a8bf791dc1f95929eb98dd9b4b8387db5f4ded557656bd090cf79c9804ef3df6

Observation dd977d04-9f5f-4d6f-82bb-f33cbf209a1f · outbound

This paper cites MiniLMv2: Multi-Head Self-Attention Relation Distillation for Compressing Pretrained Transformers.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs MiniLMv2: Multi-Head Self-Attention Relation Distillation for Compressing Pretrained Transformers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:02.299564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:02.299564Z digest=sha256:2770df29c7e1d4d1b6f3fa1e4c071b5d30e12862312b2d564044dd1441eb472a

Observation eca52e40-6877-4180-8d03-2d94c40d9d87 · outbound

This paper cites Ugc-videocaptioner: An omni ugc video detail caption model and new benchmarks, 2025.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Ugc-videocaptioner: An omni ugc video detail caption model and new benchmarks, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:02.389016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:02.389016Z digest=sha256:9f0db10a10dc3725b3eda8415a282b1072d7fe4684d8ad7e0cf76cb240382158

Observation 0bdd1548-18b8-44bd-aac7-5f0dfd162eb0 · outbound

This paper cites Qwen2.5-Omni Technical Report.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Qwen2.5-Omni Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:02.484164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:02.484164Z digest=sha256:a5ec549c940b705365916e37d37a0c1c864f8f39a9812555b5d3d9ec3b5994ae

Observation 2019e343-a879-4850-b5a1-02cdacced65f · outbound

This paper cites Qwen3-Omni Technical Report.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Qwen3-Omni Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:02.555150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:02.555150Z digest=sha256:a5ad1c08a68a1955ba15968bde4b9f5aca0a413430e18c60f8b9c0f590da7bb1

Observation 38782b91-8639-4079-acab-3b1d2adedec2 · outbound

This paper cites Avqa: A dataset for audio-visual question answering on videos.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Avqa: A dataset for audio-visual question answering on videos

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:02.635589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:02.635589Z digest=sha256:4594ac0f77feef8877329c25a7ec9110d427fa20fa14ea58af4f529b79dd9bdd

Observation 876e2cf3-fa49-41bb-b316-52fe10104723 · outbound

This paper cites Audio-centric video understanding benchmark without text shortcut.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Audio-centric video understanding benchmark without text shortcut

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:02.702999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:02.702999Z digest=sha256:74e86b019e155bdc06b3c8a1729bf4b66255694e70cf65ed4d580a180b424f15

Observation 2ee5adee-8905-4378-8267-8fc9fe604b3b · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:02.762702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:02.762702Z digest=sha256:a87fde9226950e0d81e0c9d8426940b9521f2f66d2c8896be89c8d4dede14a48

Observation f751875c-2764-451f-8691-af037b973ef2 · outbound

This paper cites MIRAI: Evaluating LLM Agents for Event Forecasting.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs MIRAI: Evaluating LLM Agents for Event Forecasting

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:02.869167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:02.869167Z digest=sha256:1c2582b97e84b038bccf763d75ce74c1ed8b846e1962a48b5e02034d6e7d00a1

Observation 9423e84a-2c86-49ad-8a40-bf144231d01e · outbound

This paper cites FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:02.970530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:02.970530Z digest=sha256:54549949eb6ea124f6201a05eea569d73e71c6a1076382cc9881dd559e33337f

Observation 8d1f9cab-1f93-4e66-859d-6073e64b8a0b · outbound

This paper cites Videollama 3: Frontier multimodal foundation models for image and video understanding, 2025.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Videollama 3: Frontier multimodal foundation models for image and video understanding, 2025

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:03.061661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:03.061661Z digest=sha256:996de18905ffab1c2cacd06bb7a0cdd4fca9fba9b204b68ba010910f84e5840c

Observation f3cb9f55-8729-4bff-ab24-21142c73b5af · outbound

This paper cites Video-LLaMA: An instruction-tuned audio-visual language model for video understanding.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Video-LLaMA: An instruction-tuned audio-visual language model for video understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:03.129471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:03.129471Z digest=sha256:1ca03ce3e9d0a3c7502f2f9f08149a9e4aa77ccbb71bc91acc906129af7c292a

Observation a7b9e03d-5771-4a71-8096-eb351e13d139 · outbound

This paper cites Timelens: Rethinking video temporal grounding with multimodal llms, 2025.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Timelens: Rethinking video temporal grounding with multimodal llms, 2025

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:03.249768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:03.249768Z digest=sha256:18e7ad9f7e9ad23c5086fc409d83b0a512e58afe0e190cd4e74761d5dcf08979

Observation 973440a7-ee75-4f4c-8cf5-03b5e5a9206f · outbound

This paper cites Llava-next: A strong zero-shot video understanding model, April 2024.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Llava-next: A strong zero-shot video understanding model, April 2024

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:03.362310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:03.362310Z digest=sha256:135622777d6953fe5d402c933ae249cf6d7acb9f2e062092ca0b352dbbae8fd5

Observation fe91e45c-5946-485a-b0a7-7b14d573777b · outbound

This paper cites Llamafactory: Unified efficient fine-tuning of 100+ language models, 2024.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Llamafactory: Unified efficient fine-tuning of 100+ language models, 2024

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:03.464504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:03.464504Z digest=sha256:f12ecdf6837b89f0d15608a05c98e580326e59947b3f0af548f8207e7a98e1cb

Observation bac7dbdc-5309-4629-b6f8-72d54bc6a17a · outbound

This paper cites Mlvu: Benchmarkingmulti-tasklongvideounderstanding.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Mlvu: Benchmarkingmulti-tasklongvideounderstanding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:03.552949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:03.552949Z digest=sha256:183c86cb2473323673e0c960edd57403d10834b5b030746a4b2e344752924e29

Observation 411ac1a0-8bad-4adb-baa0-c8db2c5edcc9 · outbound

This paper cites audio event.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs audio event

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:03.797705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:03.797705Z digest=sha256:40960d5eaa098e5e319ecd758e3103cf23eb28ec43461f2d5e36c595cbe289cd

Observation d6536786-ded1-44ac-b2ed-eaf005685b9c · outbound

This paper cites John walks down the street.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs John walks down the street

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:03.944481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:03.944481Z digest=sha256:4bbeee1d37fe4a4d5d33dbb956b4f69d61d16a0d38795daaf16ef4f2029afae0

Observation 47f4dab8-a9ee-4cfc-b335-d57bf47d48d8 · outbound

This paper cites Given the premise event: ’[Description of Event A]’, which event is its most direct conclusion?.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Given the premise event: ’[Description of Event A]’, which event is its most direct conclusion?

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:04.076476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:04.076476Z digest=sha256:d5446d380afd4c97609bfca81b613f256f49b6fbf349a477b06c6a8a49d210d8

Observation 97194122-7bbb-47b0-8e19-82ca4f0399aa · outbound

This paper cites decisive.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs decisive

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:04.254688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:04.254688Z digest=sha256:d9f62c32dc0fc92d8a394f4780f784bd785ab90a175e48dfcd60e6958ae67316

Observation 94e7e43c-22c1-4f7c-a636-ce6f24ad63b6 · outbound

This paper cites motor jam.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs motor jam

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:04.367487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:04.367487Z digest=sha256:ba7f22dfa3c19b8cc6cef0cabbebf875b3418752f01fda993ef59316f678cb13

Observation c83a88fa-3c1b-4262-9fb4-257c23e80f06 · outbound

This paper cites Understand the sequence of actions, the characters involved, and the overall flow of the narrative.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Understand the sequence of actions, the characters involved, and the overall flow of the narrative

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:04.584169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:04.584169Z digest=sha256:750d82a546d5a4ae80def86c37d9ec0b84c7ae4df4b42c75e8333d0a00237205

Observation f24d402a-5741-431c-b21d-a165a6ef7e5d · outbound

This paper cites A glass shatters.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs A glass shatters

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:04.746985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:04.746985Z digest=sha256:0da105e2910c40c3beb88a62b9dfaf9f34e899b42c8525313624ffc099937ca1

Observation 6b1280f2-bee5-4c1a-9065-57c7b474c1c3 · outbound

This paper cites an unresolved cited work.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:04.831134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:04.831134Z digest=sha256:d6565b58a56bada09452f09602a948833896110b75cb2b94e521b85ebe8936a2

Observation 374b1e23-db9e-4304-abd7-475dd096059c · outbound

This paper cites an unresolved cited work.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:04.939745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:04.939745Z digest=sha256:c97ad66e71ce06d6b2324bfb2efa33f725fe8adee502a3d2e59e18979327b933

Observation 252dd4f4-8327-4892-99ab-1869e9f6dd30 · outbound

This paper cites an unresolved cited work.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:04.987666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:04.987666Z digest=sha256:d6dfba92868e4f1e4055c0144da33f5a2eae4b198f1bae060e46496296d3c618

Observation 4ed0cc97-7cea-48bf-9c8e-1d34e1590372 · outbound

This paper cites an unresolved cited work.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Unresolved cited work

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:05.076729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:05.076729Z digest=sha256:cc2600748878c6866917b53b9a550bba0f23ebf2090edb803552f9479d184dd1

Observation 99c47ed7-c774-4274-b8eb-bd6c73c7f0cc · outbound

This paper cites # Classification Criteria: You must classify the error into exactly one of the following four categories.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs # Classification Criteria: You must classify the error into exactly one of the following four categories

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:05.164587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:05.164587Z digest=sha256:3655e1d349cc9cfa5dd2d4b93cc24a9954c9bab06f36704056364756d3995404

Observation 985fd226-c960-46c2-bf9b-ce2eb2964b74 · outbound

This paper cites ## Indicator: The visual context alone is insufficient or misleading, and the model failed because it missed the acoustic cue (e.g., a doorbell ringing off-screen).

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs ## Indicator: The visual context alone is insufficient or misleading, and the model failed because it missed the acoustic cue (e.g., a doorbell ringing off-screen)

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:05.277824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:05.277824Z digest=sha256:514aaf14aec385f5784e6081478507d1ef9f0f28e832241125bd96f02d546d6f

Observation f0b144f5-0f0e-4be2-a8cb-0c2c1884b761 · outbound

This paper cites ## Indicator: The model’s prediction contradicts clear visual evidence (e.g., predicting "driving" when the car is visually parked).

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs ## Indicator: The model’s prediction contradicts clear visual evidence (e.g., predicting "driving" when the car is visually parked)

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:05.469030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:05.469030Z digest=sha256:dd2daec407da6ece76c99f651cf7a86ce215034ed3c8b7dc8481028dd326bb20

Observation 8f7bd563-303b-4f74-b37f-b65a2ac44f91 · outbound

This paper cites ## Indicator: Understanding the scene requires prior knowledge (e.g., knowing that mixing specific chemicals causes an explosion, or knowing the rules of Chess).

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs ## Indicator: Understanding the scene requires prior knowledge (e.g., knowing that mixing specific chemicals causes an explosion, or knowing the rules of Chess)

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:05.674423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:05.674423Z digest=sha256:54470d42045b66388d3119a5769d79c1ea61a6d9d985b70600a0d5ee17bb3c59

Observation 69094d7c-2ee4-463a-853d-2df860ad8dfb · outbound

This paper cites he is exercising.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs he is exercising

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-03T09:30:05.747416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:30:05.747416Z digest=sha256:b73cbc86da43672441c3cca5415fd626349863b6a2db5c01957cdc1a66a6af97

Observation 1f515224-ed55-450d-a504-edbc430ba84e · outbound

This paper cites an unresolved cited work.

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs Unresolved cited work

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T09:29:59.991540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:29:59.991540Z digest=sha256:db36512f8a2d5d79f8256724606874838cd4340e6d40d986f01a694d67f0979d

Pith citing papers

Observation 075050cf-aab1-451b-8dc5-3b65d322515b · inbound

MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction cites this paper.

MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-06-19T17:10:48.660379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T09:26:00.413651Z digest=sha256:ae61bbef02c64048a6605c8a110a5cc108e8763207d2bd0e893d48f4ff616f23

Observation 9828d907-7b75-4569-a70c-691be18ed1f5 · inbound

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos cites this paper.

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-19T17:10:48.660379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T01:53:01.939765Z digest=sha256:8eae3a61047f6d59e7d8d4b9d4a67829d5bd9dc96260a6bf0bdff181f5b17599