Pith. sign in

Paper Citation Record · LEDGER

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation

As of 7 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 1 inbound Pith citation observation for arXiv:2507.22886.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22886 v2

Coverage vector

measured 86 of 86 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:16:10.475961Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:35:37.095627Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T10:11:08.927215Z

Reference resolution

86 of 86 outbound references displayed

  • verified exact0
  • verified fuzzy71
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1ce8bf37-9ae5-49f7-b3e4-78a3253ac6ea · outbound

This paper cites Qwen Technical Report.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.019397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.019397Z digest=sha256:a71b1d6a2b3f686897b3e55cddf8ee06fb6aaf7caa05941ae76ad2525af10ffb

Observation 2f00d2a4-c04b-4525-9a71-c51d555eb531 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.094700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.094700Z digest=sha256:3b936114258e2878c7e6a6b6b6740cb9b9caa1d3403dffec771be30a1b7e8533

Observation a5525f58-491f-4f40-93f7-0b2218537b1b · outbound

This paper cites One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.158565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.158565Z digest=sha256:8e8d1bfded2ea8ea122350e96fe2e7cc070f6b311b7f46fd894046989a44bc04

Observation 022d88cb-68fc-4714-ac2f-f3c71fb547e8 · outbound

This paper cites METEOR: An Au- tomatic Metric for MT Evaluation with Improved Correla- tion with Human Judgments.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation METEOR: An Au- tomatic Metric for MT Evaluation with Improved Correla- tion with Human Judgments

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.221846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.221846Z digest=sha256:b6cf68890d42fc0133d00edbc3ead0179a9c94bf8f87d26a6a6c4aa14868ff38

Observation 777041a2-00ee-4a1a-9dc6-3c07deffeaeb · outbound

This paper cites End-to-End Referring Video Object Segmentation with Mul- timodal Transformers.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation End-to-End Referring Video Object Segmentation with Mul- timodal Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.274355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.274355Z digest=sha256:f9103deb8cb9be12c5d732dda7994dab90e07da3cc8c43d56a1d1ca82089f8dc

Observation e8e3eb20-6961-4062-a322-8376d74a004c · outbound

This paper cites Auditory Scene Analysis: The Perceptual Organization of Sound.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Auditory Scene Analysis: The Perceptual Organization of Sound

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.346121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.346121Z digest=sha256:dc0d9a9416c4f54fe4ceafa004172b1824acf4c773ab028e100c2f9d5257083d

Observation 9784a756-a69b-4dc8-8bc6-cbd6e11ed1b5 · outbound

This paper cites COCO- Stuff: Thing and Stuff Classes in Context.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation COCO- Stuff: Thing and Stuff Classes in Context

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.422372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.422372Z digest=sha256:7003a40305af5f149df11803d1634b6e7c33d75d05562d6eff8fe9b1f633a942

Observation e53caa16-5819-4119-b0bd-d8051b52809f · outbound

This paper cites TIM: A Time Interval Ma- chine for Audio-Visual Action Recognition.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation TIM: A Time Interval Ma- chine for Audio-Visual Action Recognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.269880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:06.485367Z digest=sha256:eaf22ab7734f7aca1c3c581d1b3c763fea9fd12b286aa009ed3915da1c3b5554

Observation c505fc3d-73be-4808-9367-b907a3ce0dc7 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.252941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:06.596129Z digest=sha256:ff843e51f8ffc81eb8a9e9c80c70a63bf2ea95217b39f3ede5027496ff0a9dbc

Observation 1ae1bae4-45c0-424a-8d2f-ee1031106421 · outbound

This paper cites VGGSound: A Large-scale Audio-Visual Dataset.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation VGGSound: A Large-scale Audio-Visual Dataset

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.232394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:06.636040Z digest=sha256:b8c5cb8bbbb64703eda8977e08493ac75a2a01cf7fc46cc403afa75b251d37ac

Observation 042946cb-fc4c-4e03-a6e0-73db51bb935a · outbound

This paper cites Detect What You Can: Detecting and Representing Objects using Holistic Models and Body Parts.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Detect What You Can: Detecting and Representing Objects using Holistic Models and Body Parts

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.210671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:06.696935Z digest=sha256:a9d83b8f7de549c93e9fe5d11ec76cb9a2fffb0e58baa29e1c9d6a65db96425a

Observation 646c8340-d9cb-4d95-8eeb-802e3e114d5e · outbound

This paper cites Vision Transformer Adapter for Dense Predictions.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Vision Transformer Adapter for Dense Predictions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.190592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:06.774702Z digest=sha256:93159acf3dd663e1de4f1684fe4f582060de2d4ac7c00316db33cdff78770542

Observation 76f515e9-2c19-402f-a1c4-383e0f99abb1 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.856376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.856376Z digest=sha256:744945c1835212af5f6e30656a6356b864d8411a22422bbe3abbb97adf24d914

Observation 1e824bbf-bee8-4cc9-9ad0-a5e34c2278c8 · outbound

This paper cites Masked-attention Mask Transformer for Universal Image Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Masked-attention Mask Transformer for Universal Image Segmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.171004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:06.974243Z digest=sha256:2faaa0ce1c1ba7ab348979070df342c0fcf9942540c2c584be4880866fb91dd2

Observation e2e4208a-25f8-4a3c-bd42-3fac05f3a85d · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:07.045271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:07.045271Z digest=sha256:09caab5b006e7220763e024c5290b65bc4418ab686cafe6b4cfdd4a880fc418b

Observation b10fd106-834c-4e6d-8b7b-ad7c8fdc633b · outbound

This paper cites Qwen2-Audio Technical Report.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Qwen2-Audio Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:07.101084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:07.101084Z digest=sha256:1d95d64de773b92cf7ee900e46116fe71d7d6002de0e9834d4886d551edbeaa5

Observation f6258804-80df-43ff-9a14-46e88b13db52 · outbound

This paper cites MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.149630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:07.286988Z digest=sha256:8d1299ec0e1eaae823f655fe958e4e09513650252669bc0f2833145dc1cf9bb4

Observation 01f7c39d-536f-4319-b801-04d6ef0370b8 · outbound

This paper cites MOSE: A New Dataset for Video Object Segmentation in Complex Scenes.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MOSE: A New Dataset for Video Object Segmentation in Complex Scenes

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.128980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:07.373963Z digest=sha256:3e08dc82b39b98fa5679aff58af8cf0c6ad8975cbfa4d7f25e0dc65a851e3add

Observation 4f1abf28-9b01-48f0-80fc-2db69cc2c347 · outbound

This paper cites Multimodal referring segmentation: A survey.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Multimodal referring segmentation: A survey

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.110276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:07.468288Z digest=sha256:018d634bf9c4b8e74cfc1c14b10502f4817a60e0d5b17a970c87c86da56608bb

Observation 6c58ef9c-39ad-43f1-91f7-91a7b233dcde · outbound

This paper cites MOSEv2: A more challenging dataset for video object segmentation in complex scenes.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MOSEv2: A more challenging dataset for video object segmentation in complex scenes

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.089024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:07.555289Z digest=sha256:d7249baf9016b1dd2218d0263eb6ec8582ccad46ed7e84b151568c3ee52cda59

Observation a47ca59b-9d8f-4b35-bf3b-54d1f65feaf9 · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.070253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:07.614968Z digest=sha256:8af6813a868bb8b171a6d8d4b9effa3b321669179884a0c6bc19c429db5f2d35

Observation 48a5fa52-7c80-4b18-8b04-ac3a66fa6e90 · outbound

This paper cites A VSegFormer: Audio-Visual Segmentation with Transformer.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation A VSegFormer: Audio-Visual Segmentation with Transformer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.052365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:07.702971Z digest=sha256:38c02945f9cafb4997e1e3cf6a1ca85c5452d300c2e6acee24b8f51c5038d2bf

Observation 4dd43122-d005-4916-bbd3-36838b48a4d8 · outbound

This paper cites https://github.com/RVC- Boss/ GPT-SoVITS, 2024.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation https://github.com/RVC- Boss/ GPT-SoVITS, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.032935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:07.890555Z digest=sha256:bcbcac462a8ec7ff5df34fffbb9f86654e800f31ae66ecfe86a03c3a654e3a49

Observation 405c891e-2575-4736-bdea-30422574e22c · outbound

This paper cites Open- V ocabulary Audio-Visual Semantic Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Open- V ocabulary Audio-Visual Semantic Segmentation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.016780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:07.963845Z digest=sha256:42e9d2e778bf1ad2f1b61484194e960eebd4db2af1abb4e25889da813053b5ee

Observation 2807b96b-cd94-4504-a8de-d82e0c18893f · outbound

This paper cites Multi- modal Instruction Tuned LLMs with Fine-grained Visual Perception.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Multi- modal Instruction Tuned LLMs with Fine-grained Visual Perception

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.996594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:08.053343Z digest=sha256:c0bbb7b6075d120334de1c3ddba33434268d7e7271647bd42c60285d148ad1a5

Observation 1371932c-7134-4480-abe1-c9c9ad800365 · outbound

This paper cites Decoupling Static and Hierarchical Motion Perception for Referring Video Seg- mentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Decoupling Static and Hierarchical Motion Perception for Referring Video Seg- mentation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.977821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:08.146958Z digest=sha256:5371595b6b1f8e262ae16ebf2117dd3f6389ed3e8f1c85d355507010e5735955

Observation f6c70001-9a2d-426c-ad4f-06a8182e72a6 · outbound

This paper cites A Generalized Framework for Video Instance Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation A Generalized Framework for Video Instance Segmentation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.959155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:08.229361Z digest=sha256:2a66fb984d16551047de890bfdb7339576de6d7edce96c5206ea2b1e902dad24

Observation cd2dc9d9-8f31-40c9-a458-3c397f3283c0 · outbound

This paper cites Deep clustering: Discriminative embeddings for segmentation and separation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Deep clustering: Discriminative embeddings for segmentation and separation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.942633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:08.279887Z digest=sha256:9c87faa6f89d311334d6d844ffa793c1a2ab219fe0ea854565843689885f9d7f

Observation db461c65-20e1-4067-a559-4ffe04388662 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation LoRA: Low-Rank Adaptation of Large Language Models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.922008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:08.397805Z digest=sha256:8ad9487013263a5f8c1427eaa541766eea9971c79d8b59d9c413524ad17a4ec2

Observation d1af3745-296f-45f2-8d5f-9cc30daa77ac · outbound

This paper cites Egocentric Audio-Visual Object Localization.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Egocentric Audio-Visual Object Localization

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.905791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:08.609708Z digest=sha256:2e2e9ff10e61bdd37ed461c4b3923df7d4b48db2dc1ee0e868810115d5d40e7e

Observation 50a7f8a0-b7c0-4f1b-9ca8-a92f7895044c · outbound

This paper cites Video Object Segmentation with Language Referring Expressions.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Video Object Segmentation with Language Referring Expressions

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.888913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:08.667631Z digest=sha256:ff016934ac57ed4c8b8380613481cb4305a425a5050cdeac48c0b9e8fc05663b

Observation 672b7acd-ee7b-4197-8a0b-e39f025dd6dd · outbound

This paper cites Segment Anything.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Segment Anything

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.869618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:08.752455Z digest=sha256:56022761ce8c8f66c5d066ca67a0c933d68c67c06c6cf76d6720a41f5ab12388

Observation e8dc9a84-77ca-4455-8bb8-8e76388a25f0 · outbound

This paper cites LISA: Reasoning Segmen- tation via Large Language Model.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation LISA: Reasoning Segmen- tation via Large Language Model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.853787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:08.844260Z digest=sha256:90b0ae79e3b68be4f775706658224e9e0590ddd92a6282da2344e9aac508051b

Observation 64d9a684-9cda-4be5-9d2a-d1b0722f00db · outbound

This paper cites TVQA: Localized, Compositional Video Question Answer- ing.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation TVQA: Localized, Compositional Video Question Answer- ing

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.837725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:08.989571Z digest=sha256:fe5417e01ac80b7930e0a78496900207a5b1b7b6cc07f658b3a4f9c90bbcecdb

Observation f176f74b-fda8-43ac-901b-0ec06487a5e0 · outbound

This paper cites Learning to Answer Questions in Dynamic Audio-Visual Scenarios.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Learning to Answer Questions in Dynamic Audio-Visual Scenarios

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.819872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:09.133635Z digest=sha256:72e5bbf6944dacc09a89ffca7055df312a9f88f709a4ba8411bb5e5f913b36b2

Observation f8a8afd9-e3f0-4389-9471-37c213797fff · outbound

This paper cites Boosting Audio Visual Question Answering via Key Semantic-Aware Cues.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Boosting Audio Visual Question Answering via Key Semantic-Aware Cues

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.800716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:09.261515Z digest=sha256:528e642f41aa22ab2c175db1b3297e3eba32527a7ede54466fa3f3cf933c0be2

Observation fcccefce-30db-4420-a1eb-d1a5452bf3e4 · outbound

This paper cites Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? arXiv, 2025.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? arXiv, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.782524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:09.423301Z digest=sha256:e896c551e8f6457780065fe052f453ff938ef6b454278472ee54e8e01f8e23f1

Observation 645602e0-6408-4407-97dc-4093b00fe8fb · outbound

This paper cites Robust Referring Video Object Segmentation with Cyclic Structural Consensus.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Robust Referring Video Object Segmentation with Cyclic Structural Consensus

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.764123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:09.565012Z digest=sha256:fa5e24fb8114480dbd9d79457c8d5176f7d85d5ad22c279821f93135c1a52832

Observation c8b02ec2-3a3d-4696-8be3-73d21c9be3ea · outbound

This paper cites Baichuan-Omni Technical Report.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Baichuan-Omni Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:09.667789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:09.667789Z digest=sha256:6df77a51d76ebcf890b6e3812eab393227edc32dc4f8b33dd126664155dc357b

Observation d5aae2dc-9bb8-440b-9bc7-5ff61ddce81d · outbound

This paper cites Losh: Long-short text joint prediction network for referring video object segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Losh: Long-short text joint prediction network for referring video object segmentation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.743164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:09.765828Z digest=sha256:ee5f578309b90f0cb833e83598f9de45956f6e05960ed0d6cd41404fd7948169

Observation 50923f62-0614-498b-9011-7e18564c6710 · outbound

This paper cites GRES: Generalized Referring Expression Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation GRES: Generalized Referring Expression Segmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.725625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:09.891803Z digest=sha256:e4832134f801f3d113bfd46b071c4a26ef79b7fe079c76033f5b112349c6576b

Observation 891931f1-a7d3-4608-bc87-9c6f7b102264 · outbound

This paper cites Primitivenet: decomposing the global constraints for referring segmenta- tion.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Primitivenet: decomposing the global constraints for referring segmenta- tion

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.705305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:09.973694Z digest=sha256:d63a6e2afb572a6e15a1585fc0419cc2d1dd67d04a17d7af488b66ef4a4fda31

Observation bdfc6c6f-d5ba-4b18-93c0-8dfbcddcee92 · outbound

This paper cites Visual Instruction Tuning.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Visual Instruction Tuning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.688159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.155107Z digest=sha256:60d693f72277d9ed58b0cc967f3ea0c92bfbfaceb1602d257045d7d259f12a32

Observation 1590031a-698e-41ca-b72f-fa229f64fbbc · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Improved Baselines with Visual Instruction Tuning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.672272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.204533Z digest=sha256:caa0f672d040f565c1513d886ddf057d1207e4f9bb0cddeef1268ccebc8f0e9c

Observation 10ab099c-43fb-47a7-a078-5903e288deae · outbound

This paper cites ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Ablation Capability for Large Vision-Language Models.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Ablation Capability for Large Vision-Language Models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.654463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.210803Z digest=sha256:9eed62c5dc6ded5d7e26075c03dc04b486343de59adca498e5c63f55a9f9514d

Observation 8abb241e-01d6-4491-9e08-383ddf8654f1 · outbound

This paper cites Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Sep- aration.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Sep- aration

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.635223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.217190Z digest=sha256:f1d45bd726ce553259223f45c571de93e6e629b7c716b1ca36d7d54caf0a1f1f

Observation 015e1808-5f00-4085-98c3-22a9a3013810 · outbound

This paper cites Stepping Stones: A Progressive Training Strategy for Audio- Visual Semantic Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Stepping Stones: A Progressive Training Strategy for Audio- Visual Semantic Segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.610896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.222759Z digest=sha256:cbcb4dd40c862f1e3fdae3f6920355fcd5ea3a768f9c9b75fa277baf9931107c

Observation 9a25fee9-c7ce-46fd-a42f-0e562430d94a · outbound

This paper cites Generation and Comprehension of Unambiguous Object Descriptions.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Generation and Comprehension of Unambiguous Object Descriptions

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.588517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.228825Z digest=sha256:3bbb066141c15b5d5054f60a23b486dbf7fdffdb416a052eb8ec6ca534941788

Observation 0501c317-ebcf-452a-b993-8c1f8f8b4d2c · outbound

This paper cites V-Net: Fully Convolutional Neural Networks for V olumetric Medical Image Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation V-Net: Fully Convolutional Neural Networks for V olumetric Medical Image Segmentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.570411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.234547Z digest=sha256:481cc5b83889b0700418e39ff321c7146dc92f4110baa3768134e42d1a8ddbad

Observation 6cbd765e-9f01-4245-a1ce-0c1c298c72f3 · outbound

This paper cites https://platform.openai.com/docs/ guides/text-to-speech, 2023.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation https://platform.openai.com/docs/ guides/text-to-speech, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.552583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.247239Z digest=sha256:3d0589df26d68a76b329cebb493a041614e40c9b790097f14aeb13d77aab6d14

Observation 82d8109d-93ff-4236-b5d7-e40bc2dcf588 · outbound

This paper cites https://openai.com/index/hello- gpt-4o, 2024.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation https://openai.com/index/hello- gpt-4o, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.534588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.252563Z digest=sha256:10a5c7d31cbe1f320a4654a5ee8cf0cb4b071102bff054f87d60c7cf3e81c714

Observation 6a185d90-b484-4179-baa4-a6d681f6e7b2 · outbound

This paper cites Wnet: Audio-Guided Video Object Segmentation via Wavelet-Based Cross-Modal Denoising Networks.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Wnet: Audio-Guided Video Object Segmentation via Wavelet-Based Cross-Modal Denoising Networks

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.516684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.260334Z digest=sha256:2eae466f544fd4a7437db0f4a88c8a65926240ca91c93b99bdf753dfc9b97639

Observation 1b467605-e0aa-40c9-bc9c-023ef0f3f312 · outbound

This paper cites DetGPT: Detect What You Need via Reasoning.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation DetGPT: Detect What You Need via Reasoning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.498976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.269021Z digest=sha256:3a71316e2c959cf6e226ebf7d4dd3fc6b1bb180c5a0b862479796ac6e09359b8

Observation cbcbf9cf-d2f6-449e-bbdc-68cfc1239e0c · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.480708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.278238Z digest=sha256:f41882bfef95c290c0ba5309e24debb3da1868d14e94477b27a960702c3b39bc

Observation c563c37b-4b53-45a7-be18-589aece25f6f · outbound

This paper cites PACO: Parts and Attributes of Common Objects.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation PACO: Parts and Attributes of Common Objects

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.444625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.285742Z digest=sha256:5e03bfc918cf01176b72d566754b6a6b05171abb8db6462e27299c09b2686a7b

Observation 6e4bfd0c-8c92-4dff-ba5d-033188d7a457 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation SAM 2: Segment Anything in Images and Videos

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:10.292503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:10.292503Z digest=sha256:e6d4e6f1b4723df510661311c9bbe28faac1f9fa9b43dada78c0196f483df501

Observation 00c96345-f537-42d2-996d-8d48ef230711 · outbound

This paper cites URVOS: Unified Referring Video Object Segmentation Network with a Large-Scale Benchmark.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation URVOS: Unified Referring Video Object Segmentation Network with a Large-Scale Benchmark

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.405620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.298834Z digest=sha256:2000c7d63e046abd570787b2e0b8d23a24a50310503b1b425d9e000903002a60

Observation ec9d5370-ca5d-47a2-a75d-17e77bc45517 · outbound

This paper cites video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.364909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.304319Z digest=sha256:8ad2e2e50b946e799e2a58a9737da5e037cbc88d5f00a50fa44472f45d2b530b

Observation 43f66742-91dd-46ec-8af0-56562b948142 · outbound

This paper cites Auto- ACD: A Large-scale Dataset for Audio-Language Represen- tation Learning.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Auto- ACD: A Large-scale Dataset for Audio-Language Represen- tation Learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.347843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.312650Z digest=sha256:8504fad410186f7f0c8d088e3374dd5e8ae2a2ed5bb515e31271b3448be368b9

Observation e9f020c3-53c3-4b7e-9352-6abb95a70ee2 · outbound

This paper cites Unveiling and Mitigating Bias in Audio Visual Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Unveiling and Mitigating Bias in Audio Visual Segmentation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.330373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.319486Z digest=sha256:1f73858d707dce64d67f759a618ac9014c7636b4c910dd8790dce88e8578d18c

Observation 10441679-27b3-40d4-ae73-4bb067d5a7d8 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.311783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.324756Z digest=sha256:a58c90d37d653fdcb91b2f2de9ae5af2f3918197375e62c2ebcce33941889a3c

Observation d62a2be0-8151-4f94-8b91-b354b07c882c · outbound

This paper cites Audio-Visual Event Localization in Unconstrained Videos.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Audio-Visual Event Localization in Unconstrained Videos

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.290558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.330368Z digest=sha256:395735e6b848b5ad8e9592ba843323a10e865311ffcc59026a14e125d47f8e61

Observation 2d7d6602-7a14-4756-bee0-279c1d915fb5 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:10.336496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:10.336496Z digest=sha256:5cc456e1d68dfc7664aaa0da5a2abf6162ab829d710265097e32100265eb1270

Observation 9dc3b0b3-a9e6-4331-bf5a-a7eaa1d6b5cb · outbound

This paper cites Audio-Visual Grounding Referring Ex- pression for Robotic Manipulation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Audio-Visual Grounding Referring Ex- pression for Robotic Manipulation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.271327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.341868Z digest=sha256:d866efd7883b3ec7b26e8583614f3d1a6e9e2de4ea4345371355d27496552117

Observation 244a95f7-f82a-40e5-9e11-73152644fc18 · outbound

This paper cites Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.253156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.357613Z digest=sha256:37f5777aaffafabc499dbdd7f88b581347359b9bcdb4c555c22d1e116932bfb8

Observation 977b99c5-f402-4b6c-a997-dce4e653e001 · outbound

This paper cites Can Textual Semantics Mitigate Sounding Object Segmentation Preference? In Eur.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Can Textual Semantics Mitigate Sounding Object Segmentation Preference? In Eur

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.234846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.363089Z digest=sha256:b1f7ae68497bb62183ff9bd4c97abd861cb074a01729ace38e9ebbef0a667f03

Observation 631e99e6-604d-4d89-ad36-31efa91f39f7 · outbound

This paper cites Ref-A VS: Refer and Segment Objects in Audio-Visual Scenes.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Ref-A VS: Refer and Segment Objects in Audio-Visual Scenes

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.218188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.369265Z digest=sha256:fe1f9b2c1d7815ce1f9f0335745274084f27d43cad798b3133bde039174bd6fc

Observation 04473527-a83e-40a6-811e-0b61852742e8 · outbound

This paper cites Language as Queries for Referring Video Object Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Language as Queries for Referring Video Object Segmentation

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.198186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.374194Z digest=sha256:c0f1b44288fc83fb48212abd29b307a0290ed1428063f0c9f34cc20c7b25b81e

Observation 777ad773-26d4-453f-8f81-38b7847d8705 · outbound

This paper cites VISA: Reasoning Video Object Segmentation via Large Language Models.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.174877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.378825Z digest=sha256:ee98f8ea050bb530f19d9126efb92a903f55ec18a1562fb3d344533e435d579c

Observation 8116845e-6924-4180-b5be-c901cef114e4 · outbound

This paper cites Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.146807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.387330Z digest=sha256:75921330989ab1a5a659d0a20c5da02922023ca37c027ac14575b3ef2b9e1820

Observation 46080188-6cf1-486a-8db9-15479764098e · outbound

This paper cites A VQA: A Dataset for Audio-Visual Question Answering on Videos.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation A VQA: A Dataset for Audio-Visual Question Answering on Videos

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.121804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.394479Z digest=sha256:950c7c06e68822f49340e8648eef4f2fe61ff60cc2670a9171b3433ee17af0ff

Observation d1978e52-dd29-4c52-b280-80f4a309769b · outbound

This paper cites LA VT: Language-Aware Vision Transformer for Referring Image Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation LA VT: Language-Aware Vision Transformer for Referring Image Segmentation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.095686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.399323Z digest=sha256:42b21276377433643922245c653e127e2e01f71b0b0448658683211f995ad93a

Observation 95031c31-a3e3-417d-93f3-0937d2a67d77 · outbound

This paper cites Isda: Position-aware instance segmentation with deformable attention.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Isda: Position-aware instance segmentation with deformable attention

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.064986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.409682Z digest=sha256:3ece1ca30c9ba22b0fba2ddb3bca92f697b239978b056f1c109f77acc6e754e3

Observation 585b8fae-a251-4094-b1d6-05e743ac25d6 · outbound

This paper cites CTVIS: Consistent Training for Online Video Instance Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation CTVIS: Consistent Training for Online Video Instance Segmentation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.044795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.414531Z digest=sha256:059d132eaf6eab4c3d8baf9750040aaa8a09fa4685dfacbb4e565e5b68117b3c

Observation 4478e543-28aa-49f9-8f7f-66ef2796a7c2 · outbound

This paper cites MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.017438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.420373Z digest=sha256:f9461c6e92a29f306898744b2a977ff88beba770f63c6c8733ab4a0a52cd2203

Observation 4abbf53c-1812-4652-8e2e-20332cfcfee0 · outbound

This paper cites MOVE: Motion-guided few-shot video object segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MOVE: Motion-guided few-shot video object segmentation

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.995052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.425968Z digest=sha256:9e8188211d928e26da62dd3d50b87e3230bb4b86b33643ce9deb5f31e78b89cc

Observation e239833f-6bb0-4853-bbd8-114595adcdeb · outbound

This paper cites Modeling Context in Referring Expressions.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Modeling Context in Referring Expressions

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.964609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.430262Z digest=sha256:affba71652da646754b5d453d4dd2c90cf8dd02eb1b459389e63bf53129075ff

Observation 3f0d94cf-bbe8-4685-8d81-ca5875f7e677 · outbound

This paper cites MOTR: End-to-End Multiple-Object Tracking with Transformer.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MOTR: End-to-End Multiple-Object Tracking with Transformer

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.945246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.435156Z digest=sha256:3f6873346395360b7c99aa061a163c160de5bca2ea5dd1682eca3eed072b391e

Observation ae1f22c0-d178-4c3b-ad7c-2f9fcd04f24e · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.921392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.439995Z digest=sha256:c8b92c8add1641449302472e5d12c76f1443883025769bd2ab9ffbfee355c40c

Observation 60db074e-da48-4dee-ac7e-11891f91c7ef · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Large Language Models with Zero- initialized Attention.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation LLaMA-Adapter: Efficient Fine-tuning of Large Language Models with Zero- initialized Attention

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.902176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.444754Z digest=sha256:532faba1c420983c7f6d035a3223f230af6773a5d26fd78b00f857b5d1f0dcd6

Observation a24455e6-9922-4a95-9d94-7a0c5ac81fa1 · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:10.449340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:10.449340Z digest=sha256:95835df1256c03432b6db5e9cdccdadcd02fab0cc8b5c9210cf77b41629ede88

Observation 3dd6e515-9ec8-43cd-987f-fa1c1d08a13f · outbound

This paper cites DVIS: Decoupled Video Instance Segmentation Framework.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation DVIS: Decoupled Video Instance Segmentation Framework

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.885373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.454902Z digest=sha256:285d947afa324ea22a50da27d6ecfa3422bd6c84a55b4115a6af85d00fcd8f0c

Observation f35ef5b4-7b2b-4406-9312-0d74b021a834 · outbound

This paper cites ViLLa: Video Reasoning Segmentation with Large Language Model.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:10.460552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:10.460552Z digest=sha256:31cfb040ba876d4e91e05a12c1f83cc4ebfc33931b6f374af0d5bcc888d5d00a

Observation 77a47579-e561-4dec-99a8-67aeadf588f5 · outbound

This paper cites Scene Parsing through ADE20K Dataset.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Scene Parsing through ADE20K Dataset

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.868533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.465447Z digest=sha256:4aeef1f080271e6e0a2c390f6541661f4c13953f781a860deefefea6f20df4b4

Observation b4440455-5ce6-4e3d-bcb1-fdbf7fc961df · outbound

This paper cites Audio-Visual Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Audio-Visual Segmentation

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.852538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.470620Z digest=sha256:c70a196e93a6728e6591c11caced482ff95cf6454635c53e5d759fb0affb5aee

Observation 68aa03c1-fccc-480d-9709-0703fb587fc1 · outbound

This paper cites Tracking with Human-Intent Reasoning.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Tracking with Human-Intent Reasoning

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.832815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:16:10.475961Z digest=sha256:4cbe0f29d052e56da722c1466897ed4eec62cbd579b92fc7e45e924c89bcd8f6

Pith citing papers

Observation f4afe86d-fbd3-4785-a992-0802dc92b695 · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation

Reference 210

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:08.958133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:06959624867db697171073509a0e0770f377952d5cf8ae66e3c5047b4b4ed106