Pith. sign in

Paper Citation Record · LEDGER

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs

As of 6 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 0 inbound Pith citation observations for arXiv:2606.07643.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.07643 v1

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T14:36:53.295540Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

85 of 85 outbound references displayed

  • verified exact6
  • verified fuzzy0
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch32

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7eab8998-10dd-4032-bfff-a4ef379e2b8b · outbound

This paper cites 2024 , eprint=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs 2024 , eprint=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:438ef4b6e24561b6f1cc369c38812ff24259821f28d2eeb1d522224d28feb1b6

Observation 77902575-2ec8-45be-a020-5b7d0116048f · outbound

This paper cites GPT-4o System Card.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs GPT-4o System Card

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.430514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:84a51ddf1f78eba1beda6de6e46ad06ef406e5b1088bb9ca3948f054e3db9e9d

Observation 6eca5690-5363-4edb-b11f-58962318387d · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.434049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:238aac3fb2df1bafd533d19da6eef2875e96d366bf5e5c7c904237c98a03162f

Observation 512b2825-40ee-47e2-8b50-3a8e3177d431 · outbound

This paper cites 2023 , eprint=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs 2023 , eprint=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:0b53f724af4bd2fcfe66886d0c544be3160d5aa085b341a4f215a3ae0deb4762

Observation 6ddadbb4-f963-48a1-88e7-609000220dfe · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.420090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:874a35cc1f1f8787692ced51efe9a8d5e2e9fe34fa42cc7c484df684c01d144d

Observation 259a1f8d-3c64-4e9b-ab5d-6a18e128e250 · outbound

This paper cites The Llama 3 Herd of Models.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs The Llama 3 Herd of Models

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.423543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:ffa357c891d6d5dda3a3508022cf7a2dffde96f5c4fa27792a84f730cfcc97b5

Observation 0373cec7-b233-4249-afd1-d8bb5fb0423b · outbound

This paper cites 2023 , eprint=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs 2023 , eprint=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:df04f5270896c3154eef2a7bee17a248251c995d5068584cf407af6209e42462

Observation 5ff05621-1995-4090-a9ae-a8381507d4c7 · outbound

This paper cites Qwen2.5 Technical Report.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Qwen2.5 Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:06:21.427147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:0813b5038626a2a2587f12327449c999cf75118fb905fcfb40a627610752618c

Observation 124a9a04-d6ff-41b8-a25c-4938003809d6 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.444792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:1e52f2c64a5752ba1276179f115780b4038a71c94b4122643e4a25be7eff1bb7

Observation 1c9f38ba-f619-4fe0-b813-dcd4f1b0e1ab · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.437716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:ab1b7ebf6d9bb17a0a23c10d4dda6a3fa74a5632c662f895bb18c1b3a5bdb5d5

Observation 03cd71e0-4698-4c2e-a126-47c12f83ff40 · outbound

This paper cites Advances in neural information processing systems , volume=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Advances in neural information processing systems , volume=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:76206e8eff543fda927404bb5607b83e902e4b55703d1a115041d8f6b21aed67

Observation d9a7c800-734b-4fd5-9482-44ef9b258c56 · outbound

This paper cites 2023 , url =.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs 2023 , url =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:a1baf8cb0160b2282c89d5726c2397ab64da1981fd182a2bc172a33f4b04aafa

Observation 6b132ff5-61e0-406d-9c0f-cdad268a7692 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Advances in Neural Information Processing Systems , volume=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:3d6e4b7159290fc4e86910c0fb8d871d8c3ad997f8f13492aa1037ad6bb5dea4

Observation f73eee58-a95f-4b3b-8712-3da237ae403b · outbound

This paper cites Kimi-Audio Technical Report.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Kimi-Audio Technical Report

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.409335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:a6987478a28b00abc6e61267c8c877ffb2121afa766363e5e3250fee7b66c896

Observation 69f2883e-3ec2-4f81-9b62-78aa3388b00d · outbound

This paper cites Qwen2-Audio Technical Report.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Qwen2-Audio Technical Report

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.402186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:9d0a96f99590bbc4f17fed79ea928f0c31d328ffcfbd6c38200db4b026de7479

Observation 9fed51af-1648-4b02-b58c-48c1fbe4630b · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.406103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:c91ca8c69899aca38fa64b7c0ae948cffab2e28da6c4effd9c68596b128c8cb9

Observation f760dea0-d8a1-4938-afc0-fb811b8ab46b · outbound

This paper cites PandaGPT: One Model To Instruction-Follow Them All.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs PandaGPT: One Model To Instruction-Follow Them All

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.394725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:f7fbfcd6a6e3b0c17f7f34be8101cff0d4da146ab2828319445d0075522171bf

Observation ca55d9fa-6838-47ad-8f95-b35888bceb2a · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Advances in Neural Information Processing Systems , volume=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:366a345524db9127d38b8010ebf1cd240ca9d87b490bf445f4a053850cfe8efd

Observation 3933720a-cdc7-4355-9816-4cfd579d92fd · outbound

This paper cites Forty-first International Conference on Machine Learning , year=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Forty-first International Conference on Machine Learning , year=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:f8352a5032c658d31e093c7fbf4e2656a007c3e6a9929aa7a66f254dbd10c729

Observation 01482320-8175-4f71-ae28-60370b3353d9 · outbound

This paper cites AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.398632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:993aa39e40e44de377c7c2d1ae6da8fa426be4a0a869f600243545899fcf3b99

Observation ada64e49-89e5-4159-b6be-efc9be3dc9cb · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.412807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:e5110665024a5d881e22f182a845c66eeda026b161794b3870c4578d0edad671

Observation 03689645-6d08-4faa-b24a-2f705339135f · outbound

This paper cites 2025 , eprint=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs 2025 , eprint=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:c99b5dc8f4de71d58f3b0b148cc9cf26fa96d6ab18919b314044118b30c95378

Observation 42dd175d-8b13-42fd-a150-27dbe2c49a27 · outbound

This paper cites AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.375774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:69bff338f5328c487ab3c75dba9035eefc1e816f37afd6e11df85ae40ff0ee8e

Observation c8ba0510-7d73-4c6c-9fb1-34ac20f19080 · outbound

This paper cites 2025 , eprint=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs 2025 , eprint=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:8ecadc45025d11fff462507af19d2560b2925bd90f1e54b5023574a0698ee3d0

Observation 3c5484df-8921-4bd1-b9eb-c1a35d82b311 · outbound

This paper cites 2023 , eprint=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs 2023 , eprint=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:3602daef6b0874d9a4ea83f121667ef3699305d17a60f6160c223e331a03193e

Observation 22da7afc-03dd-4598-a03b-a8c4a95d83f2 · outbound

This paper cites 2024 , eprint=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs 2024 , eprint=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:2cd3f1bf43691a2271a365be577b98160217f4269ead145e5868fbdfff77d0e0

Observation 2b0317bc-baa6-490e-815f-2d0073361aed · outbound

This paper cites HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.372188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:295d2a19539d26b8d47bf1f84f775220a5f85a7af5911cc4f58bf7db9553298e

Observation 8b9b6cb0-7057-4580-8937-d4b939526d4d · outbound

This paper cites Baichuan-omni technical report.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Baichuan-omni technical report

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.379582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:6d174797e595a15071d64e3a9e2160394b408ecc87e110cac5b25e87817146fa

Observation 43a93783-f378-4293-8374-2d54b854bef6 · outbound

This paper cites Qwen2.5-Omni Technical Report.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Qwen2.5-Omni Technical Report

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:06:21.387364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:a9481ee14e30ada305780705d6a0a9bde7f3b72ce21410adef6321fc587facc6

Observation d71c8850-c65d-4c82-a6fe-d097db8bfe65 · outbound

This paper cites 2024 , eprint=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs 2024 , eprint=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:a015764a66e7faede9b15cba6299e449fd9ba30a4fed2b2cef338237af47f301

Observation 2cff3767-114a-4b3a-81b2-a27fb5e2465c · outbound

This paper cites 2025 , eprint=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs 2025 , eprint=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:de0962dc91ac66ce90b051f3d8b7c6d2b58b7fa93b5844d46043091ed886f769

Observation 6b7387cf-e11a-4262-87a7-3310372b1b71 · outbound

This paper cites OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.383632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:5dd799fa28b5bbd082168fbccd683110cc9a62c0bbb821be43143b895d9acebf

Observation f45cb49a-d6f0-4171-92d7-bd7672cb6225 · outbound

This paper cites In: Proceedings of the 30th ACM International Conference on Multimedia.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs In: Proceedings of the 30th ACM International Conference on Multimedia

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-28T14:42:17.912950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:c810508d4e42b6dad365926af1e3fd1a1730e10442615f450d3e17f9f9d272ff

Observation 3efb50be-37aa-4cdf-bf77-c48eb6fc0e70 · outbound

This paper cites IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , year =.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , year =

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:389df772f18fe7aacb69a84634022a8093f3db1990bb502b88898e085c40afcc

Observation 6f764bd1-c41f-4b79-b7b9-0cde6e78e852 · outbound

This paper cites WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.364975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:fbcd2e2e91793147638ed8718043c1e9615b9250fb401e119431b7445550ce0c

Observation 73b812e0-94b1-48af-aa80-06f250226bb2 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.352180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:398c51952591fe919a963c61f8d4a7c6c9e56c75556e1432b87ea984241ab5f3

Observation fbc68ea5-6d45-4c17-b63c-ed485cd01f64 · outbound

This paper cites Journal of Artificial General Intelligence , volume=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Journal of Artificial General Intelligence , volume=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:c674324fa35700d1c61637c25ca273ced35054a89fdce3722ee281666978ca45

Observation 1c55f1c5-60de-4a54-8364-479fe1057f1d · outbound

This paper cites Nature Communications , volume=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Nature Communications , volume=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:4292d1fb8187770d7d4bd4fdbb66ffc8c64f7c59f8c4aba01791a4e086a5a50d

Observation e5fbfe49-4ac7-4f20-aca4-0629f563089f · outbound

This paper cites 2023 , publisher=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs 2023 , publisher=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:eca98d730bbe686eae98e8eef9b1c5fa41c7d8e44fcad9f628f4a820c180482a

Observation f11e1a63-ec2f-4de3-8af6-49b295106a5a · outbound

This paper cites 2007 , publisher=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs 2007 , publisher=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:a0f7f189d89cbe5fd57cd348d7541bedf7ea37300c799c59f1e09df76368eca5

Observation f248a72a-a0e0-43d1-8920-a65a8acbbbc7 · outbound

This paper cites Advances in neural information processing systems , volume=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Advances in neural information processing systems , volume=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:c65c097e48fc3c996f777ed04c48229721389805dbf96531499ac9f673eea7b0

Observation 1684874b-3080-44a1-b9e1-4cefdb1db4c3 · outbound

This paper cites Forty-first International Conference on Machine Learning , year=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Forty-first International Conference on Machine Learning , year=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:b8445d67947efca959b731680d6d2bd13b4d1a79ec86d97921d6df75bf83ec0e

Observation 405639c0-0048-44e9-a116-5cdd7a8613b7 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Training Verifiers to Solve Math Word Problems

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.355934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:e3e194df0f306c87cf09f3732027c037a3a330d39a89e17ec62654dfdd4f9423

Observation b5116ad0-73cf-4d52-80f8-dd8d3cbf9026 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Evaluating Large Language Models Trained on Code

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.368488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:aec3f84649ad01ca62ebe759756488ba42a5e6e65845155ccf5b1bd307dd01a0

Observation 711dfbc5-a0cb-4863-a6cf-f4bc91566c03 · outbound

This paper cites Advances in neural information processing systems , volume=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Advances in neural information processing systems , volume=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:95a809066789c0950aeb9304a6a0be4e99c4d9932e0923d88387dea5bd49035b

Observation 8a77d57d-3757-46c9-a35a-d6a7eebef130 · outbound

This paper cites Proceedings of the European conference on computer vision (ECCV) , pages=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Proceedings of the European conference on computer vision (ECCV) , pages=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:9ddc546ef5b1f20eec708f394c8b3de029d9cc7a92e72f66725c19906165942d

Observation 358d3fbc-86f2-4537-ba00-a055eff396f1 · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:54224b2ac80045302b43b1af045c63483a2bf9cecaaae54a1592761c36818c4b

Observation 523bea54-7155-469c-9270-28c7a33d5377 · outbound

This paper cites European Conference on Computer Vision , pages=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs European Conference on Computer Vision , pages=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:21cccf77ecee00e61f4419fe398963e1581a103d098a1cf206cd07633c9cc8fd

Observation 285e210d-0a59-48f6-983a-e86204f4382d · outbound

This paper cites European Conference on Computer Vision , pages=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs European Conference on Computer Vision , pages=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:071e6533de2c59c38ce3299816f762f3939ab31f501658b74161275a1f93fcc4

Observation 8ec5c04d-9b69-4ecf-b772-9de5fe61245c · outbound

This paper cites Nature reviews neuroscience , volume=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Nature reviews neuroscience , volume=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:2420091fa2a62829da7d2d09c3aa8ee2d207fdd955bad431ff46b96562a7ce6d

Observation 8421f70d-ce34-4a31-a770-557c3de079e5 · outbound

This paper cites an unresolved cited work.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:06bc65a7ae0786e3853b4a66a6f63861545cad2d875e15fc2a6284a9f63d127d

Observation 95cedbdb-05c3-4a06-ac54-ae88dc719f89 · outbound

This paper cites an unresolved cited work.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:1a21f2dff178792ba690e80f4a5d697acf98d30546c2a0997685dde4d5c8c664

Observation 306d2ccf-2767-4b63-b476-2b7cb3111e3a · outbound

This paper cites Annual review of vision science , volume=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Annual review of vision science , volume=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:ee928864a59347a00c98ab7ca2891aa78c2dcf4b2c4ca6d14d6975416e77b665

Observation 6bc095df-d393-4fff-bd6d-09a67dc7b56d · outbound

This paper cites Current Biology , volume=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Current Biology , volume=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:dc0de797c172346b3e2e5016cc1bd17660749761fa3733edaa0b4d2e1e6cfb6f

Observation d1430315-2e72-409f-aaf1-7cc05310b29b · outbound

This paper cites Neuropsychopharmacology , volume=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Neuropsychopharmacology , volume=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:8451a443ad7798f88dde2b035a9329281d3ab96755f3dccc1053362f58291924

Observation c8b9f8af-727b-4968-881a-cdfe97acfe78 · outbound

This paper cites IEEE Open Journal of Signal Processing , year=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs IEEE Open Journal of Signal Processing , year=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:32bc129b13854ec408b89789284c8751c996ba92f9783ca75d92d8b15ed7036a

Observation ead9513f-0f4b-4682-8143-5ea973647017 · outbound

This paper cites IEEE Transactions on Pattern Analysis and Machine Intelligence , year=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs IEEE Transactions on Pattern Analysis and Machine Intelligence , year=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:5a581a937da71e25bf592c90a231c77061eaa7e7e08a41043fa3078ec359068d

Observation 65eca983-78dc-43d8-af42-ee4c9ea5a895 · outbound

This paper cites Ola: Pushing the Frontiers of Omni-Modal Language Model.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Ola: Pushing the Frontiers of Omni-Modal Language Model

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.391014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:32b671eaa3600cca6c2fdc656873f97695c910b44e68d89c836f9d5bd579bc96

Observation 8a4f4a3c-7a4c-44aa-ab98-d71ba68d7101 · outbound

This paper cites Baichuan-omni-1.5 technical report.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Baichuan-omni-1.5 technical report

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.359908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:811a0cf5bf9751feb8190607d94eacf17ce99f3631fb8fc6e45b5e2e28ef18a1

Observation 5553b942-08b1-4b22-a187-1ccbe864fa2a · outbound

This paper cites R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.416414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:5f372a4bf52d8ec50ac442646cccd00314c3e7e733391cfeb701f2492031a3a8

Observation af222d8e-548a-4726-bb9f-ee3dffeb813a · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 61

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.441366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:f1221adfa73c7aeca18550ed5915531401ec9a6ede87e585bede9c724bfcb854

Observation f137f6d5-c43a-4677-9922-2d3efa2a795b · outbound

This paper cites Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.347654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:ceba99d5e9f78f2b356d2d33d4e45a7f9d5a207fdac23673f6271859307d7bbb

Observation 43432d11-d636-49d1-ba2b-56cb6e0df3fb · outbound

This paper cites On Path to Multimodal Generalist: General-Level and General-Bench.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs On Path to Multimodal Generalist: General-Level and General-Bench

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.340815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:2648b7d441ef0f750ee02d15a9bf7d919b4f976ff49b01114f506126b3a6743f

Observation e4c6475b-f81d-4af8-a2e0-5d8304dea673 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:b69dbb6da5e059f46c69f9b0244d7daeb9cfbbeb386ebf3e1465f71719465f7e

Observation 7c1f08db-ba8b-464e-bc9a-54c54fd8d491 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 65

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.331415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:fac9de7b8865ba96f6ad7e929de6e95cc32d967ca2591a71414df4dbbd295b2f

Observation 28df9847-0e7c-41ca-b84d-889692bb0d2e · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 66

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.336404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:a0f4b5b88ebe267fa69078fbdd9664829fdf34ddf8c41c01759798c8d9f9f2b2

Observation 36603af1-061d-4692-b661-d9cc357fc349 · outbound

This paper cites European Conference on Computer Vision , pages=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs European Conference on Computer Vision , pages=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:701dcf120c3499a14f8a2f22490af567b244962c48858114a01b56e78b251174

Observation 73b3c104-6f09-47f0-87dc-56088d024a10 · outbound

This paper cites European Conference on Computer Vision , pages=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs European Conference on Computer Vision , pages=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:f3433a5d74d35719ef5ca7c32804d8c8baaf9464386f3f9009482b2fbba51787

Observation 7437bf96-f31c-41bf-a731-7a019c1d5e2e · outbound

This paper cites IEEE Access , volume=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs IEEE Access , volume=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:5496a0a8ea6ca31171120d6a3446c1dd331011ab38730e10da6e417f644aa51f

Observation 86ea63b5-35e0-4eeb-a983-593355edac8b · outbound

This paper cites Advances in neural information processing systems , volume=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Advances in neural information processing systems , volume=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:ae97e18a2a28b36e28cffbef047bce7eb9dd811bf391bc1a41edd84f2e04be0b

Observation 130d376f-fb46-4d91-be11-cd66874094dc · outbound

This paper cites International journal of Remote sensing , volume=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs International journal of Remote sensing , volume=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:bc59117e027e9558d87cc56969b1025c2a4627a4ce85220ff09435e7d218cae2

Observation bb3c1aa5-fba1-47b9-ade8-14da3e146d58 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 72

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:1a7c19a5f8ed6f14e2d9411bf693e76afaf06fbaf8563858ef0d611ac68b6f59

Observation 6f732fe6-8efe-4257-a26e-1e82583ca98a · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 73

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:c79d1646c03fe27857eb5cd6caed5aa1b7598dbb6a0455949423e36bc7b895fb

Observation d968893a-e41c-4071-98a4-1e07c77b1b77 · outbound

This paper cites ACM Transactions on Multimedia Computing, Communications and Applications , volume=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs ACM Transactions on Multimedia Computing, Communications and Applications , volume=

Reference 74

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:aa64b8b8860cee5ea05b8dd1af14f0dccc3302af6f4a94d0ec753ab21e62e0ee

Observation cc820081-9fd9-4594-9cd2-9174f3a85055 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 75

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:92fdf2178f8322d29a4fd27cc8dd79154e0b04816918e1ec1f62918cc9fbae3b

Observation 0f8aee72-c88f-4de0-b55b-5b06dbcc371f · outbound

This paper cites 2024 , isbn =.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs 2024 , isbn =

Reference 76

Resolution
verified exact
doi, observed 2026-06-28T14:42:17.915455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:951e529c0b41403dc00c9d10385e2a6f7731212ae5d1c0e1fe4a7cf4f99d8287

Observation 49e5515e-8e9d-4fa0-8b03-4f9ab08d8811 · outbound

This paper cites European Conference on Computer Vision , pages=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs European Conference on Computer Vision , pages=

Reference 77

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:8a981bd4fd4cd8d282f971dfcabb4d7395a73318d8f5984aa8901734e1591bff

Observation d3b95ef4-a88e-4e68-9463-6be83c61a2a2 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 78

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.318674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:9bbd136794cc75867d8d0c26f5d63072371c0e02902f3c20abb84a28e26c8d3c

Observation a7d7f17a-5f35-478f-9cd9-388ff19e134c · outbound

This paper cites Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.323631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:796f94c1fdd729cef4aeb6e4b35ef0fb24980a72c2bdcae0456a492c046db00d

Observation f129b174-1b38-463b-a386-c40262bf78de · outbound

This paper cites Tarsier: Recipes for Training and Evaluating Large Video Description Models.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Tarsier: Recipes for Training and Evaluating Large Video Description Models

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.312897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:54235e426c5dec7b16451b5e270d42e4abdcf75f734154dc89cc35564e0690ce

Observation 51799bc7-832b-4828-a5f4-fb065bbcfb93 · outbound

This paper cites Proceedings of the Computer Vision and Pattern Recognition Conference , pages=.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

Reference 81

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:6b842caf5245beaeba337a3f2d76daa39f3c751e8bc7cbc5724bbea5b5238ea6

Observation 95db0db5-b3ce-4464-b784-cd1cbc29a304 · outbound

This paper cites WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

Reference 82

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.293636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:b20eaebf7b6b7522abced4b6f0fbe9a36e9cfb86672c8ae68eeff24081af7066

Observation 477e8220-4d10-41cd-8307-17383bc18592 · outbound

This paper cites Mlvu: Benchmarking multi-task long video understanding.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Mlvu: Benchmarking multi-task long video understanding

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:21.299475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:98cc19d8a9b43891cc32231d3dbdeac4fc3085c034eb9df3dbe7ac66e34a400c

Observation c7fc47f5-d11c-4520-8194-9a3234e73e52 · outbound

This paper cites Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?

Reference 84

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.306035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:82d6a85b5a78783cb0c386e346fd75b5271fa9e47a439f8f4f382be5ad1729e3

Observation c8fbf330-f926-438c-89ce-90c093d3762d · outbound

This paper cites an unresolved cited work.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-06-28T14:36:53.295540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:61a7ab5fa4f2a0a1177f1ae21dd89ecb045d03d0a3b1d6bcc0030991a7d517a7

Pith citing papers

No inbound Pith citation observations are available.