Pith. sign in

Paper Citation Record · LEDGER

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs

As of 18 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 4 inbound Pith citation observations for arXiv:2506.20960.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20960 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:41:12.567839Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T03:12:38.454127Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T17:42:26.143560Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f225cebc-1161-439d-8716-a5d59df990eb · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:06.734883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:06.734883Z digest=sha256:f185260bcf427fba8ec25513a7f4fe468abf3f98d8c68e19d919a0136cf9bce6

Observation 5e2366f1-1b8a-42b4-9d74-84e83b6181e8 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Lawrence Zitnick, and Devi Parikh

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:16.535302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:41:06.815550Z digest=sha256:7e0fea622b5f6d93184f70a32b6bc243cc69a208cd8190b7fbd60a5b8c0910e9

Observation 1d587128-15c0-4c15-810d-3476be033fec · outbound

This paper cites From pre-training to fine-tuning: An in-depth analysis of large language models in the biomedical domain.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs From pre-training to fine-tuning: An in-depth analysis of large language models in the biomedical domain

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:16.366029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:41:06.894106Z digest=sha256:ba0d3f8740cb104b0a7b07c9314959e459e0209e3d113731ce8407c389af990a

Observation cd096af0-f108-46e6-9c71-25011fc1ec36 · outbound

This paper cites Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline, 2017.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline, 2017

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:06.985286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:06.985286Z digest=sha256:51b19c25e5981dab317b2143afc0d94c2fe678f2aff739d31ede59a6d3c7727d

Observation f0c1fecd-0d5e-46e2-96fd-1bf79aa69b5e · outbound

This paper cites CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:07.070388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:07.070388Z digest=sha256:f942b0f6898a191b9d41cbac09ee9af5a1478fdf685c8c447fe5cc5e2c5a38f6

Observation beb48705-2f65-4847-94ff-a778753b0bd1 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:07.185729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:07.185729Z digest=sha256:f1a24d1fba30a7e738685c3e642b802640d556ce734b9833801d3c9c1f6b53f1

Observation 98545770-40e3-4b0b-bc3c-91bc9998f5e0 · outbound

This paper cites EmotionLines: An Emotion Corpus of Multi-Party Conversations.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs EmotionLines: An Emotion Corpus of Multi-Party Conversations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:07.259689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:07.259689Z digest=sha256:2b68940e082f96adc558fb845db40218077daf7e9be6aab6dda51845dea8be2f

Observation b275f25c-94bb-4ada-ab5e-4bd3a85d5e37 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:07.346165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:07.346165Z digest=sha256:b91dc00d61de23b62c72f0e86244130143e21b7cab023a59e42ee05aea139a3d

Observation a666ecfd-3861-4c73-84bc-c7c195e80402 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:07.424327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:07.424327Z digest=sha256:e00716ae6500a0f818e4c09931231fbb9656230077f2f2bb234ce7d57fdf521e

Observation d1b67d40-cf4b-446a-8d72-87c15ab0214b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:07.513119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:07.513119Z digest=sha256:cb901e66fb23cd3424c62582e85cd99ca23a5aa69cdd16f488c7e8494eac6365

Observation 683b0b58-bb8b-431d-a76c-cfc12a45be5f · outbound

This paper cites Fleurs: Few-shot learning evaluation of universal representations of speech, 2022.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Fleurs: Few-shot learning evaluation of universal representations of speech, 2022

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:16.225901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:41:07.564476Z digest=sha256:fe205d37cb00533e5aeb6e3b9878632632f5d486581ca542336e9211d5cebe9f

Observation 4ab4286a-3fb9-438e-aa33-a05606cd76c0 · outbound

This paper cites Mmbench-video: A long-form multi-shot benchmark for holistic video understanding.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Mmbench-video: A long-form multi-shot benchmark for holistic video understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:07.676280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:07.676280Z digest=sha256:ba9423ffd07bc76d8679de49f601ead8e66d657cc970bb0e9e30e98de357f57d

Observation 3a5efb7c-427e-43b3-8172-acffd2c5344a · outbound

This paper cites Finevideo.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Finevideo

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:16.029585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:41:07.761731Z digest=sha256:ee2bf2f0fbb24fecc67f0653c027b7f5a02f783dd8b1080edc9d8705ede70b0a

Observation f7ebc80b-8bb9-418e-8453-3897aa102aad · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:07.861174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:07.861174Z digest=sha256:3f9966ceed03590d62f9e97e8b7a41b07048e4d77ac68cf70a03d1ccdce3a15f

Observation 92ec66cc-80e7-40c1-ab3b-b1c9d3761695 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:07.943880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:07.943880Z digest=sha256:9a3a922641362ee25fb65e0e9c0acafbadbe3878ec84729e7a07ae174798f257

Observation 7fdc8698-1508-428c-81bc-49d4698ad336 · outbound

This paper cites LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:08.047046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:08.047046Z digest=sha256:484f911037c642b3047c5a28b02e59104584aecbcc8c811beb44d462ec690a7d

Observation a23a837b-7bfe-465a-8445-9b69c65e6321 · outbound

This paper cites Gemini 2.5: Our most intelligent ai model, 2025.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Gemini 2.5: Our most intelligent ai model, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:15.895029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:41:08.151171Z digest=sha256:8ec90778199c88617332e67df1b5cb59a5e4df3a1bf441e257b241968fc79167

Observation c1095ca1-58ba-4ecf-b3c5-920ab8acce7a · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:08.251757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:08.251757Z digest=sha256:546ff52d0200754816351c3d85b0aaba5a84fa866525372264288c1fc4baabdf

Observation 4a4c75c9-2cc5-478f-98a8-8d6243394041 · outbound

This paper cites VizWiz Grand Challenge: Answering Visual Questions from Blind People.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs VizWiz Grand Challenge: Answering Visual Questions from Blind People

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:08.363669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:08.363669Z digest=sha256:d6e1303fc23a20f3319bf255597a6e14361625ed0181ba45bb2f0452a947c8bd

Observation caf43fd7-4a25-488e-aff0-47b4d871774b · outbound

This paper cites Measuring Massive Multitask Language Understanding.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Measuring Massive Multitask Language Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:08.456014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:08.456014Z digest=sha256:d021955c29949a96cddf31d699f4fe0958d888b002422421cd470423a76e1891

Observation 24130295-32dc-4437-8bd4-1a9d53dd9372 · outbound

This paper cites TED-LIUM 3: Twice as Much Data and Corpus Repartition for Experiments on Speaker Adaptation, page 198–208.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs TED-LIUM 3: Twice as Much Data and Corpus Repartition for Experiments on Speaker Adaptation, page 198–208

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:15.690697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:41:08.535756Z digest=sha256:b3648860f48a1de5e30421d5a70648d916e6f79caf303b0494cb49527a87041c

Observation a45e1a72-0485-486f-a6ef-fbca14729fec · outbound

This paper cites Worldsense: Evaluating real-world omnimodal understanding for multimodal llms, 2025.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Worldsense: Evaluating real-world omnimodal understanding for multimodal llms, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:08.604727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:08.604727Z digest=sha256:66f7358fe4438584569b5542bf65e75b9602c80544ae2e75f5ac75852e957a81

Observation 86971764-831a-49ba-b2e2-67cb9fdfa03c · outbound

This paper cites WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:08.706606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:08.706606Z digest=sha256:96f8a2f1b5828cffa04f5d6dae9bb60b566e74965086e84ef6dffb3758f48e5a

Observation 1067f61e-484e-4a6e-acb5-5e50ea6be5fc · outbound

This paper cites Weld, and Luke Zettlemoyer.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Weld, and Luke Zettlemoyer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:08.818096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:08.818096Z digest=sha256:83e91e10c73e9ad0186106295312ae52f02161806cbd95a3f8b361d741b7369b

Observation 91af4e98-0bde-44e4-9fb7-3cec99c1f56d · outbound

This paper cites an unresolved cited work.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:41:15.536276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:41:08.896095Z digest=sha256:cecd4244fee6be6b5ead04bd7fc4be2d86f20823153999c1e6b6cca0b106366c

Observation 7a62604c-a5a0-4205-8515-e12c0ed20d7c · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:15.399615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:41:09.002699Z digest=sha256:cefa98161ccecfcedc25fd4d9305ab58b8380f503f16f2554330526b7d96163e

Observation 36e74076-1271-4dd2-98aa-4daa1a72c325 · outbound

This paper cites Baichuan-omni-1.5 technical report.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Baichuan-omni-1.5 technical report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:09.124008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:09.124008Z digest=sha256:e9ff7e80efcb9d9d2b1d74bb615005d44c71d440e6f2397de421e1263f75c3ec

Observation 3b5ba30b-c55a-4e90-8ea9-5408d8b54426 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Evaluating object hallucination in large vision-language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:15.251245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:41:09.236197Z digest=sha256:08f1f7983c1b0827e6c80c08b5c98ca336a1908aaab27a6a0565f96a1110d585

Observation cdc5fcfe-06de-4bef-99d7-b1027923e643 · outbound

This paper cites Omnibench: Towards the future of universal omni-language models.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Omnibench: Towards the future of universal omni-language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:09.320715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:09.320715Z digest=sha256:659f0371bbe2bce63f79e75240699b0726d6c56745dbc3d46688e1ea80cfc041

Observation de84453d-89d0-429f-be37-bec0b7aab529 · outbound

This paper cites Omnibench: Towards the future of universal omni-language models, 2025.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Omnibench: Towards the future of universal omni-language models, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:09.397134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:09.397134Z digest=sha256:69bd34e162b275c5916003edafbaa0a40ab9f42ef184db07c3189fd5784fc11b

Observation 06bc4f0d-0227-4fc5-9358-d442d0b8ea74 · outbound

This paper cites StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:09.479989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:09.479989Z digest=sha256:294e9e5e94fc34f36cefbf5bbe51b8afac981757453cbc63ea605b66f690d3a5

Observation 885f63c8-9367-4978-9fb4-894fe5dcf70f · outbound

This paper cites Clotho-aqa: A crowdsourced dataset for audio question answering, 2022.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Clotho-aqa: A crowdsourced dataset for audio question answering, 2022

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:15.107148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:41:09.570093Z digest=sha256:19dc02da448f63609b18b54e2eb760d2c6025b2de558df86b5faaa4a11f4e281

Observation e6be41f8-518f-4f04-ab54-ebb24811f76f · outbound

This paper cites Visual Instruction Tuning.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Visual Instruction Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:09.665514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:09.665514Z digest=sha256:fce846e530d698b8a084872fdc33ecfd85873d20e1df57b26ad372a36b58abc9

Observation bb6f2337-76e9-4cfa-a6fc-2fc12be03d4a · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs MMBench: Is Your Multi-modal Model an All-around Player?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:09.774966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:09.774966Z digest=sha256:c780ad9f7a393261465cd26aa4addcef5729f3dbd24e3932cd97159ef3742e08

Observation c640d218-a471-4267-a56e-a0243a26a596 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:09.864310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:09.864310Z digest=sha256:764cc536dbd929cbb56160bde752f4c16013099bf7f3f0868246e52ff0de7627

Observation 8b404fbb-252f-44f4-b006-15416d8bec1b · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:14.948481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:41:09.951816Z digest=sha256:3250e8a8533b2bacabfbe183e1e67d76b2ba36ba4bde2174cb8ca4f1bf56d447

Observation 57c96686-2be7-4790-88ff-73314df941a9 · outbound

This paper cites Spoken question answering and speech continuation using spectrogram-powered LLM.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Spoken question answering and speech continuation using spectrogram-powered LLM

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:14.746493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:41:10.082574Z digest=sha256:60a16a29332a1a4c1db4e4bdc2ac65b6ea2add186b579c736d452a87db7f2311

Observation 633ee0bf-d760-4d6b-be59-1f7fdf72682b · outbound

This paper cites VoxCeleb: a large-scale speaker identification dataset.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs VoxCeleb: a large-scale speaker identification dataset

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:10.216134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:10.216134Z digest=sha256:25ca14bd8642da2f8281eff9e6a0f01f94e343319f501f979d49cecc566849e9

Observation 70251e21-663e-4876-8a5a-e6596219200f · outbound

This paper cites Minicpm-o 2.6: A gpt-4o-level mllm for vision, speech, and multi- modal live streaming on your phone.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Minicpm-o 2.6: A gpt-4o-level mllm for vision, speech, and multi- modal live streaming on your phone

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:14.604431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:41:10.316996Z digest=sha256:84536e72cd56d4e050f5ef869a946b0fc3bb21c8a23b404556aa00f2ecb9c97c

Observation 01e21f72-be9b-4566-8e57-04744ff39443 · outbound

This paper cites Librispeech: An asr corpus based on public domain audio books.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Librispeech: An asr corpus based on public domain audio books

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:14.410005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:41:10.447338Z digest=sha256:e1ff4dbbf5567350ad1dab30e3cfd984e81a4b0208c0ab3164dcc69b9e8da66d

Observation 88e76809-59f1-413b-accd-badad0c908f7 · outbound

This paper cites Plummer, Liwei Wang, Christopher M.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Plummer, Liwei Wang, Christopher M

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:14.264172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:41:10.542832Z digest=sha256:e3259b7f36b13d9728fca39a9a03195faaa5fcaabdf81ac55f64efa08b75cdc8

Observation f400eb72-c47b-4d0c-a95c-99896722aad0 · outbound

This paper cites MELD: A multimodal multi-party dataset for emotion recognition in conversation.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs MELD: A multimodal multi-party dataset for emotion recognition in conversation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:14.121328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:41:10.673146Z digest=sha256:3d4a288ac97dd3d7d8202fa709ab5aaa5bcd55fd4e815a3c0ef0393e3c7929c5

Observation 572dbd7c-2733-4873-b333-0f651e326e04 · outbound

This paper cites Question-Answering Dense Video Events.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Question-Answering Dense Video Events

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:10.769779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:10.769779Z digest=sha256:e4384e2dce96f04046e2d07cfe293a3b78390b143f2660fe4dcf3df69dac1c08

Observation 59b980dd-af92-4650-873a-384245114850 · outbound

This paper cites Learning transferable visual models from natural language supervision.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Learning transferable visual models from natural language supervision

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:10.938665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:10.938665Z digest=sha256:e75f3212e2377e6bcfdbf7a28e2be57b065490cf21dc5079059da36acb01e4f5

Observation 48d88fd4-04d5-4c15-8189-57fc55a9ba9d · outbound

This paper cites Towards VQA Models That Can Read.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Towards VQA Models That Can Read

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:11.063492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:11.063492Z digest=sha256:816b692a5505f271a908b7a1edb885fa1cb135493c224dc65c90e56e1144a989

Observation b9b2e1ea-2cde-4190-aaaa-c0fc00120c60 · outbound

This paper cites A precise detection method for transient micro short-circuit faults of lithium-ion batteries through signal processing.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs A precise detection method for transient micro short-circuit faults of lithium-ion batteries through signal processing

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T22:41:12.954686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:41:11.176834Z digest=sha256:bc2e133802b44a5a7376cfe20e208fcc8bf16f50b7a542e4420b427b7560e887

Observation a9893c17-2fa5-4e67-ae51-9b576e6ccc2c · outbound

This paper cites Language Models are Few-Shot Learners.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Language Models are Few-Shot Learners

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:11.279409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:11.279409Z digest=sha256:aa69e12c3c812bb2fad90d798320b1c8c2a02d9e3dd7b8c027ebd67dad1e6720

Observation 7f235de8-3e1a-48f2-994d-218f17a1f298 · outbound

This paper cites GPT-4 Technical Report.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs GPT-4 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:11.369391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:11.369391Z digest=sha256:4a4063fe823c89697351f64f7cf870966abcf569242287d78681df0872e0f394

Observation b785eea0-6ead-4bd2-9624-86372ae71d2a · outbound

This paper cites Qwen2.5 Technical Report.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Qwen2.5 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:11.475765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:11.475765Z digest=sha256:09024a7cb788027eb490f84fd4ee31b647d6c86cb8de31134b620dd2f733cd2b

Observation c07cb19b-8a30-4d36-a3b7-d64e99640254 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs CogVLM: Visual Expert for Pretrained Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:11.561442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:11.561442Z digest=sha256:0d8300c54e62ef5aaa9a38ab52ab47218adf549042aa39b00ef7fbc0ed6f967e

Observation 90f6c13b-1feb-4c9b-9e49-f33c3b804798 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:11.644932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:11.644932Z digest=sha256:b4b0cb5569540ae5cbe11db340c06fdacebd3c1dc3bd92cf2359a07eac81b373

Observation 50918665-cf97-4c2f-9f92-a8529f81cb66 · outbound

This paper cites Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:11.661900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:11.661900Z digest=sha256:9a6f5a20231cc34f6ba474e93e4d3eb3fbd237803a6793b2e71cf995a7a8b9e5

Observation 40f2cc99-5c71-49af-843d-2f38b6f1c043 · outbound

This paper cites Qwen2.5-Omni Technical Report.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Qwen2.5-Omni Technical Report

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:11.841114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:11.841114Z digest=sha256:f86ff374c4b984643dd742fbec2e261fb1862e37a4617d58a624fadfcef2639f

Observation 01de9829-d349-4f13-972c-fb4f35bb5ecf · outbound

This paper cites Air-bench: Benchmarking large audio-language models via generative comprehension, 2024.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Air-bench: Benchmarking large audio-language models via generative comprehension, 2024

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:11.967533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:11.967533Z digest=sha256:46faa1c4e738f754591b084adcda976a23556e89f7e739c5208efea3b82aeb78

Observation 2b3e6c2e-49c4-4f9b-8419-3f622fc5d2e1 · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions.TACL, 2:67–78, 2014.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions.TACL, 2:67–78, 2014

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:12.118730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:12.118730Z digest=sha256:ef775082af71837319fe7d22b63493fe5df4803d13055909f065ab85e39f2363

Observation d0bd4c52-9490-47df-89e2-74c141147d0c · outbound

This paper cites Berg, and Yuandong Tian.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Berg, and Yuandong Tian

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:13.941233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:41:12.231890Z digest=sha256:716e55d29a62678fa2fd860bde6159dc453745aad42e542398c64730c3af3f20

Observation 6222991a-cc70-4303-8a1e-e0bbd2d9d79c · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:12.314340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:12.314340Z digest=sha256:abacc9a2d2220f0b39530f1d82ef6e087db002ed137b043e9cf92a3218ae5527

Observation f47c004c-2c97-4e69-bf3a-a2ee632fd0f3 · outbound

This paper cites Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition, 2022.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition, 2022

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:12.402858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:12.402858Z digest=sha256:a2e816f2a3fc0e0959cc78e9446c0001767b45b7a531a66cf57c0aacebc1281c

Observation 21460dda-f689-4f1d-940c-7ec25356ea79 · outbound

This paper cites Zhang and M.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Zhang and M

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:13.810971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:41:12.455460Z digest=sha256:038efb3624299e6f7e3f90c3c5e37e2d02758def18003b4570b90497ca001eb0

Observation 524268b1-1fcb-4478-a98d-3003119ad138 · outbound

This paper cites Yin and Yang: Balancing and answering binary visual questions.

OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs Yin and Yang: Balancing and answering binary visual questions

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:41:13.679803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T22:41:12.567839Z digest=sha256:276fe0666e1b255d9bbbfb36ce031271d3729398ceba59abbace75a92822d3e2

Pith citing papers

Observation 2730f9b8-4354-44bb-b738-6016ba2b65de · inbound

SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models cites this paper.

SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:58.971096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:58.971096Z digest=sha256:97a56870758751333201ac96bb9b7be3a3374fa47d8b095328f48570a5aa3aee

Observation db3424f7-a96b-489d-bee4-18522e05b557 · inbound

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos cites this paper.

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:15:56.078287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T01:53:01.939765Z digest=sha256:5d51144b4203839b515e2dedde67b1e3254350dc92c5985c45c5488f4ade8096

Observation c5e4067b-2c44-4678-a1c2-06322a667f58 · inbound

Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition cites this paper.

Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T17:42:26.145281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T17:40:30.038337Z digest=sha256:db8f21f634a585daa7d2b4f175888063fc4070315a9686f362a8f789fba2e709

Observation cb12b0a5-33cd-4e82-98f0-de806c00ccc2 · inbound

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation cites this paper.

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T03:12:38.454127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:12:38.454127Z digest=sha256:65f59794f78590819bd4c6ec4e693a9cfba1bba31d6fa11bb7612cc231a4f03d