Pith. sign in

Paper Citation Record · LEDGER

Implicit Counterfactual Learning for Audio-Visual Segmentation

As of 20 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 0 inbound Pith citation observations for arXiv:2507.20740.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20740 v1

Coverage vector

measured 86 of 86 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:24:31.180804Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

86 of 86 outbound references displayed

  • verified exact6
  • verified fuzzy58
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 92232e28-b925-4df0-98d6-3cd58f3e6a75 · outbound

This paper cites Remembering the past and imagining the future: Common and distinct neural substrates during event construction and elaboration.

Implicit Counterfactual Learning for Audio-Visual Segmentation Remembering the past and imagining the future: Common and distinct neural substrates during event construction and elaboration

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:20.667712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:20.667712Z digest=sha256:aab739a4981a1732741877353ada0e39de04710651811aacd12f54eaef9910b0

Observation 9eb1ad31-b9f1-4f74-96f2-1c505fd5a6b4 · outbound

This paper cites Unsupervised Audio-Visual Segmentation with Modality Alignment.

Implicit Counterfactual Learning for Audio-Visual Segmentation Unsupervised Audio-Visual Segmentation with Modality Alignment

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:24:32.277437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:20.791836Z digest=sha256:ec797667d64f5580e22699ebf239330c1c99672a84f8a2a754aa0c9a52b46df7

Observation a94ba4dc-11e1-4712-acd0-31967afc369f · outbound

This paper cites Numerics of gram-schmidt orthogonalization.

Implicit Counterfactual Learning for Audio-Visual Segmentation Numerics of gram-schmidt orthogonalization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:20.951174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:20.951174Z digest=sha256:9255159e0568794a86b4e486a8f15269b58f6d31ee06fcacdc31e2f022135c48

Observation a04bf18d-2e62-4197-8c4a-71e71f82acf2 · outbound

This paper cites Self-projection and the brain.

Implicit Counterfactual Learning for Audio-Visual Segmentation Self-projection and the brain

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:21.092219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:21.092219Z digest=sha256:17b7acc4cecf013f46b3d0f9e38021ae183e83c0b438c6ab7dcc3472bd3cbefc

Observation d1420b7c-6981-4053-8105-cf2018875008 · outbound

This paper cites Bootstrapping Audio-Visual Segmentation by Strengthening Audio Cues.

Implicit Counterfactual Learning for Audio-Visual Segmentation Bootstrapping Audio-Visual Segmentation by Strengthening Audio Cues

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:24:32.180551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:21.302100Z digest=sha256:57600bd6b00cc88f6c841e1c0330e3f5fa0589d088fbcf17cb8c1ced4eb564f1

Observation f0f5345d-472c-470c-9d4f-604e5b30339e · outbound

This paper cites Unraveling in- stance associations: A closer look for audio-visual segmenta- tion.

Implicit Counterfactual Learning for Audio-Visual Segmentation Unraveling in- stance associations: A closer look for audio-visual segmenta- tion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:21.405095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:21.405095Z digest=sha256:d7dd855342c669b80edee3f5e410b9c4cf3ecfea86fed64b4a06562e82c17a3f

Observation 030edca1-1e34-4e8e-beaa-9300ac0f420a · outbound

This paper cites Cpm: Class-conditional prompting ma- chine for audio-visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Cpm: Class-conditional prompting ma- chine for audio-visual segmentation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:21.553958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:21.553958Z digest=sha256:c0ad6b34fe4235bf51c3c87776957fe6fe2816f72a666fa0a12e201a98114787

Observation a6ea029d-b907-43e9-9046-4109dee5f372 · outbound

This paper cites Cpm: Class-conditional prompting ma- chine for audio-visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Cpm: Class-conditional prompting ma- chine for audio-visual segmentation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:21.724840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:21.724840Z digest=sha256:f4e909ba52391d23e9fc0188f67a6622e85a3c1c001fe1975ed83f3459937863

Observation 1ce8f95b-8eaa-45d5-aea5-7067c2022032 · outbound

This paper cites C-cam: Causal cam for weakly supervised seman- tic segmentation on medical image.

Implicit Counterfactual Learning for Audio-Visual Segmentation C-cam: Causal cam for weakly supervised seman- tic segmentation on medical image

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.508793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:21.852776Z digest=sha256:5ccca6d30c209830c03b0dc8a8067186baffae7b93d75e1f6e6f93aa1aa5174c

Observation c6142001-4757-44a9-bc47-8aa03bb6c9c2 · outbound

This paper cites Asi- seg: Audio-driven surgical instrument segmentation with surgeon intention understanding.

Implicit Counterfactual Learning for Audio-Visual Segmentation Asi- seg: Audio-driven surgical instrument segmentation with surgeon intention understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.498445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:21.943556Z digest=sha256:f26bda49793b5e9e94d23577593e5d8d0eeb4b0c7789e9c54c546a0298a8570f

Observation fb8e7385-9045-4591-820f-a702f4b0b486 · outbound

This paper cites Regret and its avoidance: a neuroimaging study of choice behavior.

Implicit Counterfactual Learning for Audio-Visual Segmentation Regret and its avoidance: a neuroimaging study of choice behavior

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.487167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:22.148549Z digest=sha256:2dd626070e522c756ff5dda729f1329b39536c7180e1c8ca444785f0267f8753

Observation fbd22937-7218-4a6c-b92f-ef4c741c36b6 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

Implicit Counterfactual Learning for Audio-Visual Segmentation An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:22.313670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:22.313670Z digest=sha256:c81bf2227107bae4e8cd1501a5f5be641df5d6ae2f0e2913c7aa7cf224cb841f

Observation a1df29dc-3460-4ad4-b47e-84b0ecfdb377 · outbound

This paper cites Avsegformer: Audio-visual segmentation with trans- former.

Implicit Counterfactual Learning for Audio-Visual Segmentation Avsegformer: Audio-visual segmentation with trans- former

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.477009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:22.499286Z digest=sha256:cdf434811a645fe404bf813dc4463035d027be04225d89dbc5a80501ec4df400

Observation 55ba5b81-6af3-40a3-896e-5c9ab45c8ab2 · outbound

This paper cites Open- vocabulary audio-visual semantic segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Open- vocabulary audio-visual semantic segmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.467530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:22.665424Z digest=sha256:12ed9695eae30f4b90dfb85c1785d762c0d734137426f6b96f0a199ec7411581

Observation a40a676a-cf8a-4e3e-af99-e85fb9a9efcf · outbound

This paper cites Embodied intelligence via learning and evolution.

Implicit Counterfactual Learning for Audio-Visual Segmentation Embodied intelligence via learning and evolution

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.458399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:22.804976Z digest=sha256:38e62aa2cde0154ae162c67f28aee088d8760fe43e243a04e1190a65d8e5b08d

Observation 7d8ebeb7-ef7a-4eca-9289-3da185fc510d · outbound

This paper cites Improving audio-visual segmenta- tion with bidirectional generation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Improving audio-visual segmenta- tion with bidirectional generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.449069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:22.948652Z digest=sha256:4ceeb46c000788deafa2b13ff489695a1164f78786b92d63b4e8a21bbd261fc2

Observation 7db11bdb-d39e-4059-89d5-b0291245b940 · outbound

This paper cites Deep residual learning for image recognition.

Implicit Counterfactual Learning for Audio-Visual Segmentation Deep residual learning for image recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:23.073062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:23.073062Z digest=sha256:e7101f5bbf652df019a13e5f62a4f365a18a4dd110f181aec5ea89c42adf604a

Observation 7db395b5-bece-416d-a433-19a801035327 · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

Implicit Counterfactual Learning for Audio-Visual Segmentation Momentum contrast for unsupervised visual rep- resentation learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.433031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:23.261042Z digest=sha256:57d0127440566cfbfae39bbf6f9020adf6970a113639142c9ed2a51daf176aa7

Observation 11788e0e-e091-4466-80f6-be4069bbd29f · outbound

This paper cites Decoupling static and hier- archical motion perception for referring video segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Decoupling static and hier- archical motion perception for referring video segmentation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.424334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:23.491221Z digest=sha256:3ddc2833d12d268362d60f89771917eedaa47080a7543228b1f28151f5b56d08

Observation a084624f-8829-4e51-9103-40308de46c4f · outbound

This paper cites A general mechanism for perceptual decision-making in the human brain.

Implicit Counterfactual Learning for Audio-Visual Segmentation A general mechanism for perceptual decision-making in the human brain

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.415636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:24.452927Z digest=sha256:b2806217df7d8889c64f4744fb2337ae6594d376c4a3a48c5282f1b57ba912ba

Observation 25d03ac9-fee1-4fbe-b4e1-10adddf4bc64 · outbound

This paper cites Cnn archi- tectures for large-scale audio classification.

Implicit Counterfactual Learning for Audio-Visual Segmentation Cnn archi- tectures for large-scale audio classification

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.406901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:24.614224Z digest=sha256:548d65e5a681dbedd643ede3a7662051e1db4a952d6baba55712ebe3e459f50b

Observation 0439d45d-3416-427a-a8f1-2f5fcc1fe83e · outbound

This paper cites Aleatory and epistemic uncertainty in prob- ability elicitation with an example from hazardous waste management.

Implicit Counterfactual Learning for Audio-Visual Segmentation Aleatory and epistemic uncertainty in prob- ability elicitation with an example from hazardous waste management

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.397744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:24.755532Z digest=sha256:c8c01c2cc261cbfddc785f2d32ecf5e679baceb4028c88910f8e28956f42b4d9

Observation ad0cc980-5875-4b85-8a50-691e04dbeb66 · outbound

This paper cites Mix and local- ize: Localizing sound sources in mixtures.

Implicit Counterfactual Learning for Audio-Visual Segmentation Mix and local- ize: Localizing sound sources in mixtures

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.389427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:24.860562Z digest=sha256:d70790b37e62ed820afa15570271b31795ee2d9a0a0df6e05d8266d1643d4170

Observation 4bd25916-ba48-4257-b022-57a94eebb045 · outbound

This paper cites Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models.

Implicit Counterfactual Learning for Audio-Visual Segmentation Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.380217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:25.010580Z digest=sha256:0766bac572a95f945c5b195fd9d5290570b3369b77fbeb92f5533c6f0e125fa4

Observation 22163450-fa2b-4b2c-b5fa-5ccb7264a524 · outbound

This paper cites Discovering Sounding Objects by Audio Queries for Audio Visual Segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Discovering Sounding Objects by Audio Queries for Audio Visual Segmentation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:25.169098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:25.169098Z digest=sha256:747685466fcfd3f3a650c89d96421877235ebf35d74f3430067c6301d0ada00e

Observation 68326339-9d84-4988-b9c3-9c53df2ff659 · outbound

This paper cites Adaptive mixtures of local experts.Neu- ral computation, 3(1):79–87, 1991.

Implicit Counterfactual Learning for Audio-Visual Segmentation Adaptive mixtures of local experts.Neu- ral computation, 3(1):79–87, 1991

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.371190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:25.330912Z digest=sha256:be1efc68a79abed0f9350d0696de3cb5d5cfd5a0f018b9a618915dfe6bfc6396

Observation 6c4c9f0d-b966-4443-b3cb-2633c9986a55 · outbound

This paper cites Counterfactually augmented event matching for de-biased temporal sentence grounding.

Implicit Counterfactual Learning for Audio-Visual Segmentation Counterfactually augmented event matching for de-biased temporal sentence grounding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.362368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:25.559673Z digest=sha256:461dc947d53e195fc331c7bff86cb9ac5ff9e44c57f09e4a4a1903cde728c18e

Observation 993faee2-f328-4068-9126-e24260f76fb1 · outbound

This paper cites Learning to visually localize sound sources from mix- tures without prior source knowledge.

Implicit Counterfactual Learning for Audio-Visual Segmentation Learning to visually localize sound sources from mix- tures without prior source knowledge

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.353458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:25.728411Z digest=sha256:a102ee02a5394744c2d085b34f950e50c5a9d5e6c89ea70c7973832047975317

Observation dd417650-4dd9-4d51-b7d5-2064dcdfbd7e · outbound

This paper cites Segment Anything.

Implicit Counterfactual Learning for Audio-Visual Segmentation Segment Anything

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:25.802594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:25.802594Z digest=sha256:47683bdd4a0330ba447a16c33d67a28ff0558afaada5e17313c9e7e22580c3d1

Observation e63916f1-06da-44b1-9a0a-9fb875251c68 · outbound

This paper cites The singular value decompo- sition: Its computation and some applications.

Implicit Counterfactual Learning for Audio-Visual Segmentation The singular value decompo- sition: Its computation and some applications

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.344825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:25.806450Z digest=sha256:ffa97e2595738adb1f142bf62f2a2e5ccf37951a164e06ee698a73f9d33655aa

Observation 474a0f74-f627-4521-b5f7-23d07f5141c8 · outbound

This paper cites Improving vision and language concepts understanding with multimodal counterfactual samples.

Implicit Counterfactual Learning for Audio-Visual Segmentation Improving vision and language concepts understanding with multimodal counterfactual samples

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.335417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:25.892511Z digest=sha256:ae1d2dccf0c06020ee36464cec68720e05440a0454078f71cc4cc67f9dbd238a

Observation 408135f8-3618-45b7-bf4c-cdddf4533d7d · outbound

This paper cites Selm: Selective mechanism based audio-visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Selm: Selective mechanism based audio-visual segmentation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.325244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:26.058249Z digest=sha256:e0652b82c50ece6c3fc3ba5c002f0659fe002f7d8bff765060a29fc871702d48

Observation cb089aa7-6c78-41ad-8705-7df5c63882fe · outbound

This paper cites Catr: Combinatorial-dependence audio-queried transformer for audio-visual video segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Catr: Combinatorial-dependence audio-queried transformer for audio-visual video segmentation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.315333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:26.185331Z digest=sha256:cb4c9fcd92d8284fd0b6cf217a1d17d1dded719710ebac4ecac8cfb2064fdf6a

Observation 4bda86bf-c11d-4172-9739-69e6e3a052d8 · outbound

This paper cites Dice Loss for Data-imbalanced NLP Tasks.

Implicit Counterfactual Learning for Audio-Visual Segmentation Dice Loss for Data-imbalanced NLP Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:26.309876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:26.309876Z digest=sha256:5af343f1b855ba572173c96cb99765f1b9d3599073a98b137226b87718b7e1b1

Observation dc6ce626-8fa9-44c3-8a01-d4ab5140b521 · outbound

This paper cites Qdformer: Towards ro- bust audiovisual segmentation in complex environments with quantization-based semantic decomposition.

Implicit Counterfactual Learning for Audio-Visual Segmentation Qdformer: Towards ro- bust audiovisual segmentation in complex environments with quantization-based semantic decomposition

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.305293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:26.423838Z digest=sha256:d026cedde72e470344de1d0a2a50f8b0d905cf27447e06cc928282184d5e0a7b

Observation 07b5e024-bb4d-4dfc-8e01-9cb7c3ac55b4 · outbound

This paper cites Bavs: bootstrapping audio- visual segmentation by integrating foundation knowledge.

Implicit Counterfactual Learning for Audio-Visual Segmentation Bavs: bootstrapping audio- visual segmentation by integrating foundation knowledge

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.084022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:26.526751Z digest=sha256:9d464519fb5162f57cb5a3635c0a97c01b83f2e242387b9f7db257d436eb7dd5

Observation cf977619-d618-4910-b7d3-7d9da6e76caf · outbound

This paper cites Benchmarking au- dio visual segmentation for long-untrimmed videos.

Implicit Counterfactual Learning for Audio-Visual Segmentation Benchmarking au- dio visual segmentation for long-untrimmed videos

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:37.893219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:26.586204Z digest=sha256:c624f2dc623f73bc2b823cb3fed8e9a175870167996301359867c578c025edb1

Observation b31a92b5-bc64-4781-84c0-0900693af95b · outbound

This paper cites Pay attention to mlps.

Implicit Counterfactual Learning for Audio-Visual Segmentation Pay attention to mlps

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:37.688252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:26.693841Z digest=sha256:b86f0d87389d48439c8ce5f7883517a758b6db06646f4bd7a2970c9e92ca5016

Observation 83406f5f-7bb9-4ab9-8444-f8fed48f887d · outbound

This paper cites Audio-aware Query-enhanced Transformer for Audio-Visual Segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Audio-aware Query-enhanced Transformer for Audio-Visual Segmentation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:26.828928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:26.828928Z digest=sha256:7f4d48fac0a1fce428aaf4381ebbb5d1a087aff84c37ea084ce9781671a014f1

Observation 2c8f30a5-6a29-4d72-bdfd-1edacc700458 · outbound

This paper cites Annotation-free Audio-Visual Segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Annotation-free Audio-Visual Segmentation

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:24:31.926622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:26.936793Z digest=sha256:723ba93fce3c0b09ebba8184adaf9fdd2d75dc1c476532f1e9f17f467a2ed845

Observation 625ce05b-28ab-495a-b273-91ed366358e5 · outbound

This paper cites Audio-visual segmentation via unlabeled frame exploitation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Audio-visual segmentation via unlabeled frame exploitation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:37.518602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:27.047524Z digest=sha256:73358f0b11770834bfd5661c793ec7de92aab1e1287e91218bd446c6f8d7cb9f

Observation c7579be3-fa9c-49a3-be3a-10f786024ce9 · outbound

This paper cites Cross-modal causal relational reasoning for event-level visual question answer- ing.

Implicit Counterfactual Learning for Audio-Visual Segmentation Cross-modal causal relational reasoning for event-level visual question answer- ing

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:37.282649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:27.109504Z digest=sha256:1bf70a0f205f213df6cc1f810c1692a17f21ba40c59a4f035b3f44c51de5c512

Observation a06da19b-3e21-4ffb-b1d3-ac841e088740 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Implicit Counterfactual Learning for Audio-Visual Segmentation Swin transformer: Hierarchical vision transformer using shifted windows

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:27.197551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:27.197551Z digest=sha256:be0c7897f2102e193e03fe0377075b99dfcb6e3628f61d84422d7e32d1506f87

Observation 13a86cc8-f036-4e73-9a36-d75679152aaa · outbound

This paper cites Step- ping stones: a progressive training strategy for audio-visual semantic segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Step- ping stones: a progressive training strategy for audio-visual semantic segmentation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:37.131875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:27.320757Z digest=sha256:f3d58191973707c88a1d5608e3615689cf271813f64e3606b723549dad2dee51

Observation b64d31e0-fc4b-46aa-a4d2-719d938986bf · outbound

This paper cites T-vsl: Text-guided visual sound source localization in mixtures.

Implicit Counterfactual Learning for Audio-Visual Segmentation T-vsl: Text-guided visual sound source localization in mixtures

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.993803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:27.386920Z digest=sha256:5dd0a59438e8184e01854e521ff0c1836701cbfdeb8f488dd3d257b265228ed1

Observation 076f56ba-c832-4b60-b184-fb3d02046c3e · outbound

This paper cites Contrastive Conditional Latent Diffusion for Audio-visual Segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Contrastive Conditional Latent Diffusion for Audio-visual Segmentation

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:24:31.766124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:27.507201Z digest=sha256:f1d2ae4f80927d230351883119c5c16dd929d1d8b84277e8073aa96097df29e4

Observation 055e248f-983a-45a3-8192-908b2a38d400 · outbound

This paper cites Multimodal variational auto-encoder based audio-visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Multimodal variational auto-encoder based audio-visual segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.818435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:27.623102Z digest=sha256:94dfec7517c460e77c577dd5bd088a5096be012aea877a184fd57351f045684d

Observation 6bd68199-fb71-4bcd-95f9-262304aef98b · outbound

This paper cites Weakly-supervised audio- visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Weakly-supervised audio- visual segmentation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:27.707306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:27.707306Z digest=sha256:22f697af217c520890d34d148b8a811ce9039666a2b2a400218f6d0b7f300b7d

Observation 679f4081-6369-4b62-8236-66eb37155659 · outbound

This paper cites Audio-visual grouping net- work for sound localization from mixtures.

Implicit Counterfactual Learning for Audio-Visual Segmentation Audio-visual grouping net- work for sound localization from mixtures

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.647904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:27.831618Z digest=sha256:c127f13af2db5159aa6349ca968c53c110e049219c6271092c9672bd635e8250

Observation b1a01a9e-85cd-4fce-8f44-da6ae9d201d1 · outbound

This paper cites The book of why: the new science of cause and effect.

Implicit Counterfactual Learning for Audio-Visual Segmentation The book of why: the new science of cause and effect

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.501190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:27.932310Z digest=sha256:1607f53032f00c66caad440ebca52c013dabf881ea4f256971eadb541126ec1d

Observation bcabcb0c-4af6-402f-ac10-adf0ba432cfd · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Implicit Counterfactual Learning for Audio-Visual Segmentation Learning transferable visual models from natural language supervi- sion

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:28.047650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:28.047650Z digest=sha256:91d4c8834f5e83351955f1c4c580ce83b3bd164594828664c3cca222432c1fd7

Observation 550e563d-bb74-48bb-8da8-98f2af648330 · outbound

This paper cites Zero-shot text-to-image generation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Zero-shot text-to-image generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:28.155249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:28.155249Z digest=sha256:48d4a3ad70483f5202ff35dde0d0bbb8f9ee81145a7d7eff41abd21c6d39a35d

Observation a0617c11-9af7-4ca2-aed7-99a53b509b4b · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Implicit Counterfactual Learning for Audio-Visual Segmentation High-resolution image synthesis with latent diffusion models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.393079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:28.258287Z digest=sha256:4b3a7f252188d8aa3b2c483a27580301181a8d48abeaecd5bf6f3f4f3e3f4f7f

Observation 5e9251e1-5e22-4c0c-87d5-841fd30864bb · outbound

This paper cites Focal loss for dense ob- ject detection.

Implicit Counterfactual Learning for Audio-Visual Segmentation Focal loss for dense ob- ject detection

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:28.331897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:28.331897Z digest=sha256:3bdcae6ddd4d7c2b318d8f14225b2fdfc40325d81021c3885d2caba1afa93a5c

Observation af45f266-bf12-4f4b-8d4c-88b9a904e7be · outbound

This paper cites The wasserstein distance and approxi- mation theorems.

Implicit Counterfactual Learning for Audio-Visual Segmentation The wasserstein distance and approxi- mation theorems

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.249885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:28.439113Z digest=sha256:754500999ef23674bf997b10952aca0780c85735b7adf05bdfcf93afb51f478e

Observation f7ab43a7-705a-400e-92ea-0d19b46e369c · outbound

This paper cites Better aggregation in test-time augmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Better aggregation in test-time augmentation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.097862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:28.519677Z digest=sha256:7c903445a09fa8ca61a1386132d2553bd7d5a9f6bd5698835d9cbcbe89e8ddb6

Observation ffb19ff6-108c-4007-954d-feb8189ad7d3 · outbound

This paper cites A mathematical theory of commu- nication.

Implicit Counterfactual Learning for Audio-Visual Segmentation A mathematical theory of commu- nication

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.925906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:28.634597Z digest=sha256:052878fc16cf61f35d40619a2baa20ee39b8ee53fa1fe376ec369c2c4266afd7

Observation 99535a89-fd96-4c7c-bd56-9077b0df0bef · outbound

This paper cites Looking similar sounding different: Leveraging counterfactual cross-modal pairs for audiovisual representa- tion learning.

Implicit Counterfactual Learning for Audio-Visual Segmentation Looking similar sounding different: Leveraging counterfactual cross-modal pairs for audiovisual representa- tion learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.804483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:28.704904Z digest=sha256:d2c78d47f720e0add7080faaf7a5108624fd4d36962ca1e86a42527e15be0f04

Observation 23110967-391c-4226-986c-db595b6da29e · outbound

This paper cites 3d audio-visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation 3d audio-visual segmentation

Reference 59

Resolution
verified exact
raw_fallback, observed 2026-08-06T13:24:31.614704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:28.777772Z digest=sha256:4ffac121c63e676e1a73706c4a4cf157865c02bb41646d43ff9aeff9e68332d0

Observation e2a02759-95dd-43e7-ac93-50bfdd89d177 · outbound

This paper cites Unveiling and mitigating bias in audio visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Unveiling and mitigating bias in audio visual segmentation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.702149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:28.858392Z digest=sha256:f90d4d50f93b028dae57d0705b4478e34e90eaffec440cd25e482651092109fd

Observation 15377690-589c-47c1-8b04-87f7639ab53a · outbound

This paper cites Learning audio-visual source localization via false negative aware contrastive learning.

Implicit Counterfactual Learning for Audio-Visual Segmentation Learning audio-visual source localization via false negative aware contrastive learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.572731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:28.979464Z digest=sha256:5001ea2714912e44e08017c9d6308d5587aaaf9bc906a92b8d4553dbf9293507

Observation 2515ccf1-998f-4f4d-ac57-f3a328a0e817 · outbound

This paper cites Language-guided audio-visual source separation via trimodal consistency.

Implicit Counterfactual Learning for Audio-Visual Segmentation Language-guided audio-visual source separation via trimodal consistency

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.417838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:29.061286Z digest=sha256:0a78b761c5bd2d965d75be7ce5d379d6067f236a9fe38d59e2602e769723e35b

Observation 0c2e86cd-6ec9-4654-a9a2-c8867e76ca0a · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Implicit Counterfactual Learning for Audio-Visual Segmentation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:29.167740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:29.167740Z digest=sha256:9ebb1f90a49aab017b4f0bc993992476f51a0c78c193244a2abb99bc315c1d11

Observation 8a2ca42f-3d2f-4d93-a5d3-fb2d02c7faac · outbound

This paper cites Neural discrete representation learning.

Implicit Counterfactual Learning for Audio-Visual Segmentation Neural discrete representation learning

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.232776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:29.269444Z digest=sha256:12dde5891504dabd13f34ee20b884d17ce06ef4fdc504dba9b42663eb902c90d

Observation f2efaf7a-148a-4976-ac2b-eb2b10eaeb84 · outbound

This paper cites Counterfactual cycle-consistent learn- ing for instruction following and generation in vision- language navigation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Counterfactual cycle-consistent learn- ing for instruction following and generation in vision- language navigation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.985658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:29.340354Z digest=sha256:dab7bc9f49b975d771cb77067cac5a2490855c86a0ae2b52cf5c986b901e242a

Observation dfe75355-433d-4e55-8086-5c7bb61ea5b4 · outbound

This paper cites Vision-and-language naviga- tion via causal learning.

Implicit Counterfactual Learning for Audio-Visual Segmentation Vision-and-language naviga- tion via causal learning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.817968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:29.450485Z digest=sha256:d1113830acf775a718a86392dd6e8c5caa255607f2ab957fa154ef48de7d844d

Observation 6b0363dd-d240-4a07-b11c-99257ee012d2 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Implicit Counterfactual Learning for Audio-Visual Segmentation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:29.575827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:29.575827Z digest=sha256:d607773c8928de1797732b287d5874033b3832bb1763fce31fd9c5f0b2aaf15c

Observation 57326a31-7fed-419a-9e03-f04f782c88ca · outbound

This paper cites Pvt v2: Improved baselines with pyramid vision transformer.

Implicit Counterfactual Learning for Audio-Visual Segmentation Pvt v2: Improved baselines with pyramid vision transformer

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.591109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:29.651671Z digest=sha256:3d935d1d6df969c0d7ab63dd54bc5e8763794c4188c60f5dfb57c0f3a5238ad2

Observation 2d8e8a11-5f77-4b10-8d4c-a21c080864b3 · outbound

This paper cites Drivedreamer: Towards real-world- drive world models for autonomous driving.

Implicit Counterfactual Learning for Audio-Visual Segmentation Drivedreamer: Towards real-world- drive world models for autonomous driving

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:29.738949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:29.738949Z digest=sha256:1b247757be8022d506e2e02be75ecb809983769b60553576781cd4b3c74950d2

Observation b00d8a91-7e04-44c4-be8e-414d4a62e7f1 · outbound

This paper cites Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer.

Implicit Counterfactual Learning for Audio-Visual Segmentation Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:24:31.372411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:29.839727Z digest=sha256:1859d848a8f24dc8da651ab23491d4c621bfab646fbdd34f71b0f02fa933c3ea

Observation d84413af-577e-40c0-bc13-1a6e731270b3 · outbound

This paper cites Prompting segmentation with sound is gen- eralizable audio-visual source localizer.

Implicit Counterfactual Learning for Audio-Visual Segmentation Prompting segmentation with sound is gen- eralizable audio-visual source localizer

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.366851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:29.911767Z digest=sha256:c34fb6b6d5f136df09d1cb687b76b5c8babc5237da1a1a2afb98259fff3978cf

Observation 2e1162cf-9d17-4e68-b0cc-8e8d71258c23 · outbound

This paper cites Can textual semantics mitigate sounding object segmentation preference? In European Conference on Com- puter Vision, pages 340–356.

Implicit Counterfactual Learning for Audio-Visual Segmentation Can textual semantics mitigate sounding object segmentation preference? In European Conference on Com- puter Vision, pages 340–356

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.206530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:30.002229Z digest=sha256:db5b12073a18a8674037a7428c8c25a87cf7a6106f52db8a49670e3a3d124105

Observation 145afdec-8d79-44d4-a233-1db959720e1c · outbound

This paper cites Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.033283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:30.078314Z digest=sha256:56a03fb9c2c85c8415ddbb8fb0d1358e9c1d4369fe720266b89ddcd053d5f2d8

Observation 2f93e96f-c19c-41b4-8489-c070a09d9865 · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

Implicit Counterfactual Learning for Audio-Visual Segmentation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:30.193705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:30.193705Z digest=sha256:ea223d071eabb7197d278b2df5abcd0c79002c80a7cf3fa38fb1aea4d468af3b

Observation 25e4326e-a60c-4e3f-b7ee-99d4bb643f4a · outbound

This paper cites Referred by multi-modality: A unified tem- poral transformer for video object segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Referred by multi-modality: A unified tem- poral transformer for video object segmentation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.919685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:30.264676Z digest=sha256:5dbd5e0e4d28d226ebcb37b02311d13eaf5f61a5028594ddba5d8ace32fdd08f

Observation 5434c471-de08-4491-804c-0091f009c7c1 · outbound

This paper cites Cooperation does matter: Exploring multi-order bilateral relations for audio- visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Cooperation does matter: Exploring multi-order bilateral relations for audio- visual segmentation

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.757091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:30.395792Z digest=sha256:12ae7c4f13c437d9f6fb60f5cecda2c9df057832f8c4d2239924eed1403eb916

Observation c1f36da8-5497-4e1f-9186-b2268aa953a2 · outbound

This paper cites Revisiting counterfactual prob- lems in referring expression comprehension.

Implicit Counterfactual Learning for Audio-Visual Segmentation Revisiting counterfactual prob- lems in referring expression comprehension

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.617230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:30.482824Z digest=sha256:4a5a8a22b2728cf516826d9a8ae99c923860fb0f8b6707ce135cd981b7a059b9

Observation fc20ea18-d27a-4797-a966-e1b8bd72dcee · outbound

This paper cites Losh: Long-short text joint prediction network for referring video object segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Losh: Long-short text joint prediction network for referring video object segmentation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.459830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:30.591093Z digest=sha256:b2732fd63d16f25efdeba057d72eeb5be544faa90e233a02f591bbef89dcbd61

Observation 9c325e08-f56f-4645-b594-f7a34b8e575a · outbound

This paper cites Discovering the real association: Multimodal causal rea- soning in video question answering.

Implicit Counterfactual Learning for Audio-Visual Segmentation Discovering the real association: Multimodal causal rea- soning in video question answering

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.321674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:30.648645Z digest=sha256:bfc83a115a4f607b30e517b13fde91b1bad64f450d1d8bb5b863478b74d57ab6

Observation de25b491-8657-4d1f-8b76-893d79779aad · outbound

This paper cites Weakly- supervised mirror detection via scribble annotations.

Implicit Counterfactual Learning for Audio-Visual Segmentation Weakly- supervised mirror detection via scribble annotations

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.147692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:30.711366Z digest=sha256:137ae0fd388b6d10c69105bc949fff2383035652ca26b5e7970f603b86ce1095

Observation 7410b16a-d8ab-4b01-b633-cbee76be50da · outbound

This paper cites Heterogeneous experts and hierarchical perception for un- derwater salient object detection.

Implicit Counterfactual Learning for Audio-Visual Segmentation Heterogeneous experts and hierarchical perception for un- derwater salient object detection

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:32.969158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:30.787086Z digest=sha256:ccf948e2c641607387b4b98ce98b9b307ca29573944008af4cae6b5ae9e5fbf9

Observation 5310ff46-d21b-4fc5-836b-00592f7203dc · outbound

This paper cites Causal intervention for weakly- supervised semantic segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Causal intervention for weakly- supervised semantic segmentation

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:32.874213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:30.860023Z digest=sha256:258c0a1cac5ca55f81191b5f1418f2fdde4360982fa60e1c6cb95995bcff25cd

Observation 919094cc-387d-44d6-9e58-d36771852dd4 · outbound

This paper cites Audio–visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Audio–visual segmentation

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:30.918766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:30.918766Z digest=sha256:a1ed792de269ef249936108f0cdcaa79b308b0659bb7e90dfd6f7ab09d42ea4d

Observation ed7a1761-fd70-4dab-bcb5-a3e28308dedc · outbound

This paper cites Audio-visual segmentation with semantics.

Implicit Counterfactual Learning for Audio-Visual Segmentation Audio-visual segmentation with semantics

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:32.591154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:31.109915Z digest=sha256:75ec568e911b2df736ac0f8928ae50b72cf8d9f3891c6e68b44fe7b2163ce3d8

Observation cc711852-876c-4aff-83cc-ffd8923a2ae3 · outbound

This paper cites Exploring pre-trained text- to-video diffusion models for referring video object segmen- tation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Exploring pre-trained text- to-video diffusion models for referring video object segmen- tation

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:32.424693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:31.180804Z digest=sha256:0c313db0fb76c9e41e4935960577b55043ca64b2b6a48ed5570c9f760fecf2da

Observation 9abe49c0-d912-40e2-9f47-04af5b5e656a · outbound

This paper cites 2, 5, 6, 8.

Implicit Counterfactual Learning for Audio-Visual Segmentation 2, 5, 6, 8

Reference 403

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:32.709159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T13:24:31.029697Z digest=sha256:5acd57b4c3711e067c18b17b34ebd5c52648576d07023dd5e566c23327b6abd1

Pith citing papers

No inbound Pith citation observations are available.