Pith. sign in

Paper Citation Record · LEDGER

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation

As of 7 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 1 inbound Pith citation observation for arXiv:2507.22886.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22886 v2

Coverage vector

measured 86 of 86 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:16:10.475961Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:35:37.095627Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T10:11:08.927215Z

Reference resolution

86 of 86 outbound references displayed

  • verified exact0
  • verified fuzzy71
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1ce8bf37-9ae5-49f7-b3e4-78a3253ac6ea · outbound

This paper cites Qwen Technical Report.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.019397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.019397Z digest=sha256:a71b1d6a2b3f686897b3e55cddf8ee06fb6aaf7caa05941ae76ad2525af10ffb

Observation 2f00d2a4-c04b-4525-9a71-c51d555eb531 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.094700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.094700Z digest=sha256:3b936114258e2878c7e6a6b6b6740cb9b9caa1d3403dffec771be30a1b7e8533

Observation a5525f58-491f-4f40-93f7-0b2218537b1b · outbound

This paper cites One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.158565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.158565Z digest=sha256:8e8d1bfded2ea8ea122350e96fe2e7cc070f6b311b7f46fd894046989a44bc04

Observation 022d88cb-68fc-4714-ac2f-f3c71fb547e8 · outbound

This paper cites METEOR: An Au- tomatic Metric for MT Evaluation with Improved Correla- tion with Human Judgments.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation METEOR: An Au- tomatic Metric for MT Evaluation with Improved Correla- tion with Human Judgments

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.221846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.221846Z digest=sha256:b6cf68890d42fc0133d00edbc3ead0179a9c94bf8f87d26a6a6c4aa14868ff38

Observation 777041a2-00ee-4a1a-9dc6-3c07deffeaeb · outbound

This paper cites End-to-End Referring Video Object Segmentation with Mul- timodal Transformers.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation End-to-End Referring Video Object Segmentation with Mul- timodal Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.274355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.274355Z digest=sha256:f9103deb8cb9be12c5d732dda7994dab90e07da3cc8c43d56a1d1ca82089f8dc

Observation e8e3eb20-6961-4062-a322-8376d74a004c · outbound

This paper cites Auditory Scene Analysis: The Perceptual Organization of Sound.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Auditory Scene Analysis: The Perceptual Organization of Sound

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.346121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.346121Z digest=sha256:dc0d9a9416c4f54fe4ceafa004172b1824acf4c773ab028e100c2f9d5257083d

Observation 9784a756-a69b-4dc8-8bc6-cbd6e11ed1b5 · outbound

This paper cites COCO- Stuff: Thing and Stuff Classes in Context.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation COCO- Stuff: Thing and Stuff Classes in Context

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.422372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.422372Z digest=sha256:7003a40305af5f149df11803d1634b6e7c33d75d05562d6eff8fe9b1f633a942

Observation e53caa16-5819-4119-b0bd-d8051b52809f · outbound

This paper cites TIM: A Time Interval Ma- chine for Audio-Visual Action Recognition.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation TIM: A Time Interval Ma- chine for Audio-Visual Action Recognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.269880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:06.485367Z digest=sha256:f97ba67b5242fac3e703f34055399ee35dc4060a5b7d88f124e1aebfd9110c67

Observation c505fc3d-73be-4808-9367-b907a3ce0dc7 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.252941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:06.596129Z digest=sha256:1d6300fd5029361d7dddb174330d0c9682cfb8874baaf092862b89530dda0652

Observation 1ae1bae4-45c0-424a-8d2f-ee1031106421 · outbound

This paper cites VGGSound: A Large-scale Audio-Visual Dataset.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation VGGSound: A Large-scale Audio-Visual Dataset

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.232394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:06.636040Z digest=sha256:a830e970e386cbe38f67e4439ae9c933f2777719d1d57727ad6047f2a02498cd

Observation 042946cb-fc4c-4e03-a6e0-73db51bb935a · outbound

This paper cites Detect What You Can: Detecting and Representing Objects using Holistic Models and Body Parts.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Detect What You Can: Detecting and Representing Objects using Holistic Models and Body Parts

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.210671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:06.696935Z digest=sha256:364c3db9a314b649fd97198728d140e7d56c89ee34aff545907976c289ed35cb

Observation 646c8340-d9cb-4d95-8eeb-802e3e114d5e · outbound

This paper cites Vision Transformer Adapter for Dense Predictions.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Vision Transformer Adapter for Dense Predictions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.190592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:06.774702Z digest=sha256:8ff56223d696e1f3abdc00f28e4f8bc31225cd9e757ac277429fe395288d9489

Observation 76f515e9-2c19-402f-a1c4-383e0f99abb1 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.856376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.856376Z digest=sha256:744945c1835212af5f6e30656a6356b864d8411a22422bbe3abbb97adf24d914

Observation 1e824bbf-bee8-4cc9-9ad0-a5e34c2278c8 · outbound

This paper cites Masked-attention Mask Transformer for Universal Image Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Masked-attention Mask Transformer for Universal Image Segmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.171004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:06.974243Z digest=sha256:8db9b1001ff7b56d2fb473a06d67fb3137fe3b0c233fe1cf638f82b32441c5fd

Observation e2e4208a-25f8-4a3c-bd42-3fac05f3a85d · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:07.045271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:07.045271Z digest=sha256:09caab5b006e7220763e024c5290b65bc4418ab686cafe6b4cfdd4a880fc418b

Observation b10fd106-834c-4e6d-8b7b-ad7c8fdc633b · outbound

This paper cites Qwen2-Audio Technical Report.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Qwen2-Audio Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:07.101084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:07.101084Z digest=sha256:1d95d64de773b92cf7ee900e46116fe71d7d6002de0e9834d4886d551edbeaa5

Observation f6258804-80df-43ff-9a14-46e88b13db52 · outbound

This paper cites MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.149630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:07.286988Z digest=sha256:226276aff627f5ccc88097b1be0a194ceb9546cba7866fe5b684c751799bc0bf

Observation 01f7c39d-536f-4319-b801-04d6ef0370b8 · outbound

This paper cites MOSE: A New Dataset for Video Object Segmentation in Complex Scenes.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MOSE: A New Dataset for Video Object Segmentation in Complex Scenes

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.128980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:07.373963Z digest=sha256:05aba01101245ee02d4510db6c931754f6c65f74e88d7120c444ca34e4934881

Observation 4f1abf28-9b01-48f0-80fc-2db69cc2c347 · outbound

This paper cites Multimodal referring segmentation: A survey.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Multimodal referring segmentation: A survey

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.110276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:07.468288Z digest=sha256:b071a7d07aa9e9becc343fd150d19036327be0d2e5b22c3afb7ff95a4b963157

Observation 6c58ef9c-39ad-43f1-91f7-91a7b233dcde · outbound

This paper cites MOSEv2: A more challenging dataset for video object segmentation in complex scenes.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MOSEv2: A more challenging dataset for video object segmentation in complex scenes

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.089024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:07.555289Z digest=sha256:66de04afd14ffcf4e97ef5b0da7fdb70c3700d6b0e2b0c2028d61c6d94b13d65

Observation a47ca59b-9d8f-4b35-bf3b-54d1f65feaf9 · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.070253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:07.614968Z digest=sha256:016ad151e3e256a4c420a754edec43a824fb68018e9b5b78db3f67b9a6593a14

Observation 48a5fa52-7c80-4b18-8b04-ac3a66fa6e90 · outbound

This paper cites A VSegFormer: Audio-Visual Segmentation with Transformer.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation A VSegFormer: Audio-Visual Segmentation with Transformer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.052365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:07.702971Z digest=sha256:473a82b6cf6e1d0b43e01d97fd8f6d90975efbb7e53879ef4ff9279723d00610

Observation 4dd43122-d005-4916-bbd3-36838b48a4d8 · outbound

This paper cites https://github.com/RVC- Boss/ GPT-SoVITS, 2024.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation https://github.com/RVC- Boss/ GPT-SoVITS, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.032935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:07.890555Z digest=sha256:622a949d7fdcf14be050d5c4ff01783d822bba51488e573cac4e4e02b2d5fdcf

Observation 405c891e-2575-4736-bdea-30422574e22c · outbound

This paper cites Open- V ocabulary Audio-Visual Semantic Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Open- V ocabulary Audio-Visual Semantic Segmentation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.016780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:07.963845Z digest=sha256:699ee7b6b7d8041dc0f0388322cf55665adbbf7dedf02e3c32fd01acd3e0f7e3

Observation 2807b96b-cd94-4504-a8de-d82e0c18893f · outbound

This paper cites Multi- modal Instruction Tuned LLMs with Fine-grained Visual Perception.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Multi- modal Instruction Tuned LLMs with Fine-grained Visual Perception

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.996594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:08.053343Z digest=sha256:4cbd1bd8e2bcbf3328d81dbf8754c93a3c14c3907c49f714090f1b52072b89d9

Observation 1371932c-7134-4480-abe1-c9c9ad800365 · outbound

This paper cites Decoupling Static and Hierarchical Motion Perception for Referring Video Seg- mentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Decoupling Static and Hierarchical Motion Perception for Referring Video Seg- mentation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.977821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:08.146958Z digest=sha256:b7bf842de5c011d4bb2ef221b5ec1b24da3579e4faf70ff02dafea40f71e02d1

Observation f6c70001-9a2d-426c-ad4f-06a8182e72a6 · outbound

This paper cites A Generalized Framework for Video Instance Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation A Generalized Framework for Video Instance Segmentation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.959155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:08.229361Z digest=sha256:bdbf128a58a4bee1f9a2fa08f4c606db4f3b67319f8b8672e1f07c1342255ccb

Observation cd2dc9d9-8f31-40c9-a458-3c397f3283c0 · outbound

This paper cites Deep clustering: Discriminative embeddings for segmentation and separation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Deep clustering: Discriminative embeddings for segmentation and separation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.942633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:08.279887Z digest=sha256:944c821cc23a65827fca5c23dccf8fb2097f781f7c14d69f2b205c3aa4e2f7d8

Observation db461c65-20e1-4067-a559-4ffe04388662 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation LoRA: Low-Rank Adaptation of Large Language Models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.922008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:08.397805Z digest=sha256:88ae035d63955f2b26bf53f4c76272543ebf3c6fa14d92455425a32d8fb56e09

Observation d1af3745-296f-45f2-8d5f-9cc30daa77ac · outbound

This paper cites Egocentric Audio-Visual Object Localization.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Egocentric Audio-Visual Object Localization

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.905791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:08.609708Z digest=sha256:fd6d74569710bd223ce5c66b6a671fc48d786ec08570bae3e0b345870d5d9ee1

Observation 50a7f8a0-b7c0-4f1b-9ca8-a92f7895044c · outbound

This paper cites Video Object Segmentation with Language Referring Expressions.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Video Object Segmentation with Language Referring Expressions

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.888913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:08.667631Z digest=sha256:454103692cb4d65a4233f0e49e055d600e63149f7de279a530a932ee20c3e0e4

Observation 672b7acd-ee7b-4197-8a0b-e39f025dd6dd · outbound

This paper cites Segment Anything.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Segment Anything

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.869618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:08.752455Z digest=sha256:7ccd840db5da53fa973c1f6762369edca5d7b24b46405459cb33a06b39314a33

Observation e8dc9a84-77ca-4455-8bb8-8e76388a25f0 · outbound

This paper cites LISA: Reasoning Segmen- tation via Large Language Model.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation LISA: Reasoning Segmen- tation via Large Language Model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.853787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:08.844260Z digest=sha256:ecac3a9ae9568841ae4841b683c34fdcf3278850088adb3d0e525f10dc4dc988

Observation 64d9a684-9cda-4be5-9d2a-d1b0722f00db · outbound

This paper cites TVQA: Localized, Compositional Video Question Answer- ing.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation TVQA: Localized, Compositional Video Question Answer- ing

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.837725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:08.989571Z digest=sha256:98c08517ceb9c33a35e9fe48b2590d844f588fcd18c2f12dd0ef1d64be5ee663

Observation f176f74b-fda8-43ac-901b-0ec06487a5e0 · outbound

This paper cites Learning to Answer Questions in Dynamic Audio-Visual Scenarios.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Learning to Answer Questions in Dynamic Audio-Visual Scenarios

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.819872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:09.133635Z digest=sha256:e511e589225374be0eb2c6c7d4b8222ea64079b8a5f000851eb6dbe08b640c9e

Observation f8a8afd9-e3f0-4389-9471-37c213797fff · outbound

This paper cites Boosting Audio Visual Question Answering via Key Semantic-Aware Cues.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Boosting Audio Visual Question Answering via Key Semantic-Aware Cues

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.800716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:09.261515Z digest=sha256:8520ee1d4096dc2faefe966ef66accc50ab38316f99e6708097907840b5f373c

Observation fcccefce-30db-4420-a1eb-d1a5452bf3e4 · outbound

This paper cites Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? arXiv, 2025.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? arXiv, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.782524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:09.423301Z digest=sha256:fd43cbcdff6680b52f1afed246f22c2f78a99f6698e8847b4bd22dccdf732b2b

Observation 645602e0-6408-4407-97dc-4093b00fe8fb · outbound

This paper cites Robust Referring Video Object Segmentation with Cyclic Structural Consensus.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Robust Referring Video Object Segmentation with Cyclic Structural Consensus

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.764123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:09.565012Z digest=sha256:228c68ae0a704773244a6166a9046245cc699653cd4b9eabc6d034ee606756ec

Observation c8b02ec2-3a3d-4696-8be3-73d21c9be3ea · outbound

This paper cites Baichuan-Omni Technical Report.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Baichuan-Omni Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:09.667789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:09.667789Z digest=sha256:6df77a51d76ebcf890b6e3812eab393227edc32dc4f8b33dd126664155dc357b

Observation d5aae2dc-9bb8-440b-9bc7-5ff61ddce81d · outbound

This paper cites Losh: Long-short text joint prediction network for referring video object segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Losh: Long-short text joint prediction network for referring video object segmentation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.743164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:09.765828Z digest=sha256:01b915461c457ad6b6c4ed7465a1aca0bb07a58812881ca208871bfd377ef325

Observation 50923f62-0614-498b-9011-7e18564c6710 · outbound

This paper cites GRES: Generalized Referring Expression Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation GRES: Generalized Referring Expression Segmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.725625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:09.891803Z digest=sha256:73f03633fbdf65ed84b710b99670d5217769b625b22b27a2349c557a6c1226dc

Observation 891931f1-a7d3-4608-bc87-9c6f7b102264 · outbound

This paper cites Primitivenet: decomposing the global constraints for referring segmenta- tion.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Primitivenet: decomposing the global constraints for referring segmenta- tion

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.705305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:09.973694Z digest=sha256:bca561e6054ee4144fb0da4994f99acc651f6046df43e07d32ded4740513c22d

Observation bdfc6c6f-d5ba-4b18-93c0-8dfbcddcee92 · outbound

This paper cites Visual Instruction Tuning.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Visual Instruction Tuning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.688159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.155107Z digest=sha256:691b25b272c8df69f896ad4a2bacd5de463df9d9b5c86e6f1bf62e76c15ef820

Observation 1590031a-698e-41ca-b72f-fa229f64fbbc · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Improved Baselines with Visual Instruction Tuning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.672272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.204533Z digest=sha256:576683df67a185c33bdd2e93dcf93d4d33f3b73746d8aa71b4c3e9fbdba11901

Observation 10ab099c-43fb-47a7-a078-5903e288deae · outbound

This paper cites ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Ablation Capability for Large Vision-Language Models.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Ablation Capability for Large Vision-Language Models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.654463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.210803Z digest=sha256:b43e69ab66141454296ee237d33d131e09eefabf88616c4ecf0ea48c50a6fb2f

Observation 8abb241e-01d6-4491-9e08-383ddf8654f1 · outbound

This paper cites Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Sep- aration.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Sep- aration

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.635223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.217190Z digest=sha256:16a6ca07f7a8e09f5444f4ec8b08f0c3e89fb5ded12ac37ee28c8d7eef5124d4

Observation 015e1808-5f00-4085-98c3-22a9a3013810 · outbound

This paper cites Stepping Stones: A Progressive Training Strategy for Audio- Visual Semantic Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Stepping Stones: A Progressive Training Strategy for Audio- Visual Semantic Segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.610896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.222759Z digest=sha256:6b45a065b640a0523c79d2a985ee44c0726f5e7b9712b38f30a48d55e135d60a

Observation 9a25fee9-c7ce-46fd-a42f-0e562430d94a · outbound

This paper cites Generation and Comprehension of Unambiguous Object Descriptions.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Generation and Comprehension of Unambiguous Object Descriptions

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.588517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.228825Z digest=sha256:368061a516918dda00fa461004de231e1109d957656a2cd3314450768c057e2b

Observation 0501c317-ebcf-452a-b993-8c1f8f8b4d2c · outbound

This paper cites V-Net: Fully Convolutional Neural Networks for V olumetric Medical Image Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation V-Net: Fully Convolutional Neural Networks for V olumetric Medical Image Segmentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.570411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.234547Z digest=sha256:b5aa1664762e56842ed3bd755f3ef3e536d391d5dca4efbd2f81032a693e8c6c

Observation 6cbd765e-9f01-4245-a1ce-0c1c298c72f3 · outbound

This paper cites https://platform.openai.com/docs/ guides/text-to-speech, 2023.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation https://platform.openai.com/docs/ guides/text-to-speech, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.552583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.247239Z digest=sha256:d4e77406c44b5de00ba0e80a054525896be9a8e5b7cc267456daf1ca7921ab79

Observation 82d8109d-93ff-4236-b5d7-e40bc2dcf588 · outbound

This paper cites https://openai.com/index/hello- gpt-4o, 2024.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation https://openai.com/index/hello- gpt-4o, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.534588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.252563Z digest=sha256:63a169185adb10d8a38c1306da955ecf6a3f10e31d3494dbf735b61c4bf22ab6

Observation 6a185d90-b484-4179-baa4-a6d681f6e7b2 · outbound

This paper cites Wnet: Audio-Guided Video Object Segmentation via Wavelet-Based Cross-Modal Denoising Networks.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Wnet: Audio-Guided Video Object Segmentation via Wavelet-Based Cross-Modal Denoising Networks

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.516684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.260334Z digest=sha256:474857cac05507d2ddb449e4058d6a810b6186a813b34ecd32c839b714f9fab6

Observation 1b467605-e0aa-40c9-bc9c-023ef0f3f312 · outbound

This paper cites DetGPT: Detect What You Need via Reasoning.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation DetGPT: Detect What You Need via Reasoning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.498976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.269021Z digest=sha256:d06e132dff9172e1eb7f82ee1fc9c88aea06262c544303b10b322a649c04130e

Observation cbcbf9cf-d2f6-449e-bbdc-68cfc1239e0c · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.480708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.278238Z digest=sha256:698788b94ea50cd2794171a442042976fb527781f70e668c6ee6a555041ae17a

Observation c563c37b-4b53-45a7-be18-589aece25f6f · outbound

This paper cites PACO: Parts and Attributes of Common Objects.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation PACO: Parts and Attributes of Common Objects

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.444625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.285742Z digest=sha256:db0b01e1365175ce6111c9045243c25d32a94b395795bb513f836ba4926e95f8

Observation 6e4bfd0c-8c92-4dff-ba5d-033188d7a457 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation SAM 2: Segment Anything in Images and Videos

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:10.292503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:10.292503Z digest=sha256:e6d4e6f1b4723df510661311c9bbe28faac1f9fa9b43dada78c0196f483df501

Observation 00c96345-f537-42d2-996d-8d48ef230711 · outbound

This paper cites URVOS: Unified Referring Video Object Segmentation Network with a Large-Scale Benchmark.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation URVOS: Unified Referring Video Object Segmentation Network with a Large-Scale Benchmark

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.405620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.298834Z digest=sha256:18dc574b77a2ecc65251d8a8fd76da00711465aeb17a1dccab3d16ef8a0d2ad7

Observation ec9d5370-ca5d-47a2-a75d-17e77bc45517 · outbound

This paper cites video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.364909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.304319Z digest=sha256:986763bd2b8e352462a3f8c075b12f119bb8ba279ab38e5397115ebadb06ac82

Observation 43f66742-91dd-46ec-8af0-56562b948142 · outbound

This paper cites Auto- ACD: A Large-scale Dataset for Audio-Language Represen- tation Learning.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Auto- ACD: A Large-scale Dataset for Audio-Language Represen- tation Learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.347843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.312650Z digest=sha256:994676bb7ef6708156ba6bea098adb447eaa6cf4ca6c5d4d67b5d9b90558cca2

Observation e9f020c3-53c3-4b7e-9352-6abb95a70ee2 · outbound

This paper cites Unveiling and Mitigating Bias in Audio Visual Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Unveiling and Mitigating Bias in Audio Visual Segmentation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.330373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.319486Z digest=sha256:5fde03f0ad237aac70153f294afe61fa5279a258cc8fb7014e3eebd28e3789ae

Observation 10441679-27b3-40d4-ae73-4bb067d5a7d8 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.311783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.324756Z digest=sha256:e72bd0c691b8a00da9ef2657e7f15b1da31a7fed9dba40abfa193dd6dbd092de

Observation d62a2be0-8151-4f94-8b91-b354b07c882c · outbound

This paper cites Audio-Visual Event Localization in Unconstrained Videos.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Audio-Visual Event Localization in Unconstrained Videos

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.290558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.330368Z digest=sha256:657c11effe99699d6e25559153c273b1d0b445b48f3a8587d96a8f3e2ed732e5

Observation 2d7d6602-7a14-4756-bee0-279c1d915fb5 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:10.336496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:10.336496Z digest=sha256:5cc456e1d68dfc7664aaa0da5a2abf6162ab829d710265097e32100265eb1270

Observation 9dc3b0b3-a9e6-4331-bf5a-a7eaa1d6b5cb · outbound

This paper cites Audio-Visual Grounding Referring Ex- pression for Robotic Manipulation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Audio-Visual Grounding Referring Ex- pression for Robotic Manipulation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.271327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.341868Z digest=sha256:3f3b67eec421cb44e0a8bf0a1ce053bb2147ae6d1226ecb1fc82d9a513f4d9d2

Observation 244a95f7-f82a-40e5-9e11-73152644fc18 · outbound

This paper cites Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.253156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.357613Z digest=sha256:09e15a3c4bbc9f635526e4a2e0e1d1a07c083ba6cbf308e92217da93d252f207

Observation 977b99c5-f402-4b6c-a997-dce4e653e001 · outbound

This paper cites Can Textual Semantics Mitigate Sounding Object Segmentation Preference? In Eur.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Can Textual Semantics Mitigate Sounding Object Segmentation Preference? In Eur

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.234846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.363089Z digest=sha256:9774eddc077748617bb32a3f93d39c66d186c43828be3657075ec5cc0cd413a6

Observation 631e99e6-604d-4d89-ad36-31efa91f39f7 · outbound

This paper cites Ref-A VS: Refer and Segment Objects in Audio-Visual Scenes.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Ref-A VS: Refer and Segment Objects in Audio-Visual Scenes

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.218188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.369265Z digest=sha256:b973790f7146b42353a4c03965513e3d3da9782fe5cfe22f9cd493575b74385a

Observation 04473527-a83e-40a6-811e-0b61852742e8 · outbound

This paper cites Language as Queries for Referring Video Object Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Language as Queries for Referring Video Object Segmentation

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.198186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.374194Z digest=sha256:2bd0165fd619906fdffbd1927b6caec2df91871654112fc518f7c4c4a7690252

Observation 777ad773-26d4-453f-8f81-38b7847d8705 · outbound

This paper cites VISA: Reasoning Video Object Segmentation via Large Language Models.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.174877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.378825Z digest=sha256:eaa41f85a9f082177552c2ecc55900d25702a10357733e480f1acfe718291f1b

Observation 8116845e-6924-4180-b5be-c901cef114e4 · outbound

This paper cites Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.146807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.387330Z digest=sha256:ded42f00429bcc777ab384f05ea78fd285c481ef586d2667e04303ae904fe3ed

Observation 46080188-6cf1-486a-8db9-15479764098e · outbound

This paper cites A VQA: A Dataset for Audio-Visual Question Answering on Videos.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation A VQA: A Dataset for Audio-Visual Question Answering on Videos

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.121804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.394479Z digest=sha256:222077c473dff180448905fb34f5e8d5169556e7b2784d3e56b187b8845182ac

Observation d1978e52-dd29-4c52-b280-80f4a309769b · outbound

This paper cites LA VT: Language-Aware Vision Transformer for Referring Image Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation LA VT: Language-Aware Vision Transformer for Referring Image Segmentation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.095686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.399323Z digest=sha256:eb6f17c3f93209b0ef316b3fa38cfdb5037782003e402b34c607be342b29b28d

Observation 95031c31-a3e3-417d-93f3-0937d2a67d77 · outbound

This paper cites Isda: Position-aware instance segmentation with deformable attention.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Isda: Position-aware instance segmentation with deformable attention

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.064986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.409682Z digest=sha256:9b679ef9b01112b0c94682acce953dc03a8bcfdc9f5370441c09aa792a114dd2

Observation 585b8fae-a251-4094-b1d6-05e743ac25d6 · outbound

This paper cites CTVIS: Consistent Training for Online Video Instance Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation CTVIS: Consistent Training for Online Video Instance Segmentation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.044795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.414531Z digest=sha256:4d800837aeae6ab9f29a80d05cbbc7f99617bc08fd87dfa0855c17f71cf5b63a

Observation 4478e543-28aa-49f9-8f7f-66ef2796a7c2 · outbound

This paper cites MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.017438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.420373Z digest=sha256:39b01c645b009b650a8dcce876a39dffda145a57aa74aa13eb93a00e1c4ee0d1

Observation 4abbf53c-1812-4652-8e2e-20332cfcfee0 · outbound

This paper cites MOVE: Motion-guided few-shot video object segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MOVE: Motion-guided few-shot video object segmentation

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.995052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.425968Z digest=sha256:0aa6b6192233fc27a7237885ce9cff1ed600e62a5459384f66a2794dae4af429

Observation e239833f-6bb0-4853-bbd8-114595adcdeb · outbound

This paper cites Modeling Context in Referring Expressions.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Modeling Context in Referring Expressions

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.964609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.430262Z digest=sha256:c3128a946fae1b8174fb816bb92bb843c01a97269e6882b78cd49a10d8a407a3

Observation 3f0d94cf-bbe8-4685-8d81-ca5875f7e677 · outbound

This paper cites MOTR: End-to-End Multiple-Object Tracking with Transformer.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MOTR: End-to-End Multiple-Object Tracking with Transformer

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.945246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.435156Z digest=sha256:01b1f3e14220ea5567f80454dcc2e7df21ea45e392afea3020cd3632036caa65

Observation ae1f22c0-d178-4c3b-ad7c-2f9fcd04f24e · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.921392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.439995Z digest=sha256:3c0061fc775bb8007907177f3860a31fb8489e4eb0c3d91b6b0b7d4d834b1aaf

Observation 60db074e-da48-4dee-ac7e-11891f91c7ef · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Large Language Models with Zero- initialized Attention.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation LLaMA-Adapter: Efficient Fine-tuning of Large Language Models with Zero- initialized Attention

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.902176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.444754Z digest=sha256:b997b9ce0ad91e3d0c804a142c9ea9402e7d707ea6046e9294c0bf42baa83b80

Observation a24455e6-9922-4a95-9d94-7a0c5ac81fa1 · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:10.449340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:10.449340Z digest=sha256:95835df1256c03432b6db5e9cdccdadcd02fab0cc8b5c9210cf77b41629ede88

Observation 3dd6e515-9ec8-43cd-987f-fa1c1d08a13f · outbound

This paper cites DVIS: Decoupled Video Instance Segmentation Framework.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation DVIS: Decoupled Video Instance Segmentation Framework

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.885373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.454902Z digest=sha256:845852ac1aa9dfdddac85591a7a067461c045c88b095b6b53d03a20ad5f0d373

Observation f35ef5b4-7b2b-4406-9312-0d74b021a834 · outbound

This paper cites ViLLa: Video Reasoning Segmentation with Large Language Model.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:10.460552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:10.460552Z digest=sha256:31cfb040ba876d4e91e05a12c1f83cc4ebfc33931b6f374af0d5bcc888d5d00a

Observation 77a47579-e561-4dec-99a8-67aeadf588f5 · outbound

This paper cites Scene Parsing through ADE20K Dataset.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Scene Parsing through ADE20K Dataset

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.868533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.465447Z digest=sha256:0510a23c7f42dfa0ab936f944bcaeb1d93815b7f38712bef24626908ef2b19dd

Observation b4440455-5ce6-4e3d-bcb1-fdbf7fc961df · outbound

This paper cites Audio-Visual Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Audio-Visual Segmentation

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.852538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.470620Z digest=sha256:b9d208beef2bb0fc2b33a018a8ff5eb80973d89e97cfa199d84ea6bd2d288863

Observation 68aa03c1-fccc-480d-9709-0703fb587fc1 · outbound

This paper cites Tracking with Human-Intent Reasoning.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Tracking with Human-Intent Reasoning

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.832815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T11:16:10.475961Z digest=sha256:43ed01efbb2dd72511bac41a19b9dc4cf9fe6d061b06129f5dd5e3032073aec6

Pith citing papers

Observation f4afe86d-fbd3-4785-a992-0802dc92b695 · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation

Reference 210

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:08.958133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:4905a4ba2535b99fc242693857ece91d195d2b283dae449a59333a2ce3d1b45b