Pith. sign in

Paper Citation Record · LEDGER

Implicit Counterfactual Learning for Audio-Visual Segmentation

As of 7 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 0 inbound Pith citation observations for arXiv:2507.20740.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20740 v1

Coverage vector

measured 86 of 86 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:24:31.180804Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

86 of 86 outbound references displayed

  • verified exact6
  • verified fuzzy58
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 92232e28-b925-4df0-98d6-3cd58f3e6a75 · outbound

This paper cites Remembering the past and imagining the future: Common and distinct neural substrates during event construction and elaboration.

Implicit Counterfactual Learning for Audio-Visual Segmentation Remembering the past and imagining the future: Common and distinct neural substrates during event construction and elaboration

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:20.667712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:20.667712Z digest=sha256:94447719f35a41fd0b1e14dc03099dcf0e01d96ba8acab53c31594ff2c559ec2

Observation 9eb1ad31-b9f1-4f74-96f2-1c505fd5a6b4 · outbound

This paper cites Unsupervised Audio-Visual Segmentation with Modality Alignment.

Implicit Counterfactual Learning for Audio-Visual Segmentation Unsupervised Audio-Visual Segmentation with Modality Alignment

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:24:32.277437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:20.791836Z digest=sha256:a874b518f6a9d9a0d03fd42bcfd72cc6eb816db017b41f9c5ee9d0f773b5a447

Observation a94ba4dc-11e1-4712-acd0-31967afc369f · outbound

This paper cites Numerics of gram-schmidt orthogonalization.

Implicit Counterfactual Learning for Audio-Visual Segmentation Numerics of gram-schmidt orthogonalization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:20.951174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:20.951174Z digest=sha256:462143099f12c2e33ffe93fa6ecba9afcf331cd03206a2fbfa9a752d00a0fc78

Observation a04bf18d-2e62-4197-8c4a-71e71f82acf2 · outbound

This paper cites Self-projection and the brain.

Implicit Counterfactual Learning for Audio-Visual Segmentation Self-projection and the brain

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:21.092219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:21.092219Z digest=sha256:1de93d0f5327b4ff4f92bca7f11790ebc666f444ba2120dbf3a9b97cd052f303

Observation d1420b7c-6981-4053-8105-cf2018875008 · outbound

This paper cites Bootstrapping Audio-Visual Segmentation by Strengthening Audio Cues.

Implicit Counterfactual Learning for Audio-Visual Segmentation Bootstrapping Audio-Visual Segmentation by Strengthening Audio Cues

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:24:32.180551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:21.302100Z digest=sha256:8bf7ce99b2358070e0536610783c9cf91a6a48b049bb3074a1dcdec9ca10f02e

Observation f0f5345d-472c-470c-9d4f-604e5b30339e · outbound

This paper cites Unraveling in- stance associations: A closer look for audio-visual segmenta- tion.

Implicit Counterfactual Learning for Audio-Visual Segmentation Unraveling in- stance associations: A closer look for audio-visual segmenta- tion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:21.405095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:21.405095Z digest=sha256:5a9f742a8a7d6a8eb1693cdebf01915d0c391f8a8452a825752b19314ee683e4

Observation 030edca1-1e34-4e8e-beaa-9300ac0f420a · outbound

This paper cites Cpm: Class-conditional prompting ma- chine for audio-visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Cpm: Class-conditional prompting ma- chine for audio-visual segmentation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:21.553958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:21.553958Z digest=sha256:7303b4809869fe3904ad655af0f445e6ad915b73cfc8df35c8a19e7bf1f31ca9

Observation a6ea029d-b907-43e9-9046-4109dee5f372 · outbound

This paper cites Cpm: Class-conditional prompting ma- chine for audio-visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Cpm: Class-conditional prompting ma- chine for audio-visual segmentation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:21.724840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:21.724840Z digest=sha256:ee0d9192f6273afbfd53c791629a30a602a629a3710c880fe1a7b9339ededbfc

Observation 1ce8f95b-8eaa-45d5-aea5-7067c2022032 · outbound

This paper cites C-cam: Causal cam for weakly supervised seman- tic segmentation on medical image.

Implicit Counterfactual Learning for Audio-Visual Segmentation C-cam: Causal cam for weakly supervised seman- tic segmentation on medical image

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.508793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:21.852776Z digest=sha256:6df33272a631d845dbde694e87542bd02916acf322a2ac2d45d2ec15e96d30a2

Observation c6142001-4757-44a9-bc47-8aa03bb6c9c2 · outbound

This paper cites Asi- seg: Audio-driven surgical instrument segmentation with surgeon intention understanding.

Implicit Counterfactual Learning for Audio-Visual Segmentation Asi- seg: Audio-driven surgical instrument segmentation with surgeon intention understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.498445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:21.943556Z digest=sha256:7dd02b1d460d84b8aa33edc14568fa1e16e0aa39af6c2ae94e51df3935db6d17

Observation fb8e7385-9045-4591-820f-a702f4b0b486 · outbound

This paper cites Regret and its avoidance: a neuroimaging study of choice behavior.

Implicit Counterfactual Learning for Audio-Visual Segmentation Regret and its avoidance: a neuroimaging study of choice behavior

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.487167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:22.148549Z digest=sha256:beb72cd1264b691fac5172144d51e092a9ff61ef26031d01dda238d1f0685235

Observation fbd22937-7218-4a6c-b92f-ef4c741c36b6 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

Implicit Counterfactual Learning for Audio-Visual Segmentation An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:22.313670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:22.313670Z digest=sha256:0d19dc81bb326b81328e517626dc1e865eaad21a3bdb3ad942080a78a062f87c

Observation a1df29dc-3460-4ad4-b47e-84b0ecfdb377 · outbound

This paper cites Avsegformer: Audio-visual segmentation with trans- former.

Implicit Counterfactual Learning for Audio-Visual Segmentation Avsegformer: Audio-visual segmentation with trans- former

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.477009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:22.499286Z digest=sha256:cac42367697d80e55f74bca887e7687e64b027a0aeb8ac0a997a77fa5c8083ce

Observation 55ba5b81-6af3-40a3-896e-5c9ab45c8ab2 · outbound

This paper cites Open- vocabulary audio-visual semantic segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Open- vocabulary audio-visual semantic segmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.467530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:22.665424Z digest=sha256:306060b3af141c837754fa8d1698001c985eafdb72c07284a3c2d0a4aebd0432

Observation a40a676a-cf8a-4e3e-af99-e85fb9a9efcf · outbound

This paper cites Embodied intelligence via learning and evolution.

Implicit Counterfactual Learning for Audio-Visual Segmentation Embodied intelligence via learning and evolution

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.458399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:22.804976Z digest=sha256:3068751ab0a4550e4cf2d664fc082100cfd689089a9af8e0cf8a2aceb312bff2

Observation 7d8ebeb7-ef7a-4eca-9289-3da185fc510d · outbound

This paper cites Improving audio-visual segmenta- tion with bidirectional generation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Improving audio-visual segmenta- tion with bidirectional generation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.449069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:22.948652Z digest=sha256:aa04a363a76c079854f984e5755dd9f3004feb3e34f1cebd42819cb038161705

Observation 7db11bdb-d39e-4059-89d5-b0291245b940 · outbound

This paper cites Deep residual learning for image recognition.

Implicit Counterfactual Learning for Audio-Visual Segmentation Deep residual learning for image recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:23.073062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:23.073062Z digest=sha256:eb2c061f3efbf3b9e5236cbf21156a4fd0cdd4e9882e74ff54823b73e7712dcb

Observation 7db395b5-bece-416d-a433-19a801035327 · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

Implicit Counterfactual Learning for Audio-Visual Segmentation Momentum contrast for unsupervised visual rep- resentation learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.433031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:23.261042Z digest=sha256:3f33b13d48322e6abef2c7e81557b0a5643d23daa38c2870da0d3e630316ee01

Observation 11788e0e-e091-4466-80f6-be4069bbd29f · outbound

This paper cites Decoupling static and hier- archical motion perception for referring video segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Decoupling static and hier- archical motion perception for referring video segmentation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.424334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:23.491221Z digest=sha256:cf0276c570224586070fb7bcd90b9f47b85644f4dde4ce3004b17f0d31bb1fbc

Observation a084624f-8829-4e51-9103-40308de46c4f · outbound

This paper cites A general mechanism for perceptual decision-making in the human brain.

Implicit Counterfactual Learning for Audio-Visual Segmentation A general mechanism for perceptual decision-making in the human brain

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.415636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:24.452927Z digest=sha256:4873e640b9fae4210cf087a833dac4789a008897bb71b902018d9e6474cc4a0a

Observation 25d03ac9-fee1-4fbe-b4e1-10adddf4bc64 · outbound

This paper cites Cnn archi- tectures for large-scale audio classification.

Implicit Counterfactual Learning for Audio-Visual Segmentation Cnn archi- tectures for large-scale audio classification

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.406901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:24.614224Z digest=sha256:f807ef1a3c7e98b2c134cbdc47e7e25e473c2690fc047543cf0eec5d094f0fe3

Observation 0439d45d-3416-427a-a8f1-2f5fcc1fe83e · outbound

This paper cites Aleatory and epistemic uncertainty in prob- ability elicitation with an example from hazardous waste management.

Implicit Counterfactual Learning for Audio-Visual Segmentation Aleatory and epistemic uncertainty in prob- ability elicitation with an example from hazardous waste management

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.397744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:24.755532Z digest=sha256:125c5acf8d5527abc6ecf00596c9816fb3f5cd5c8d2b82e7523a203692cb7771

Observation ad0cc980-5875-4b85-8a50-691e04dbeb66 · outbound

This paper cites Mix and local- ize: Localizing sound sources in mixtures.

Implicit Counterfactual Learning for Audio-Visual Segmentation Mix and local- ize: Localizing sound sources in mixtures

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.389427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:24.860562Z digest=sha256:a42088ecd3155a030a249534f0086a0656308bafe46be84d476f8cde1e43371c

Observation 4bd25916-ba48-4257-b022-57a94eebb045 · outbound

This paper cites Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models.

Implicit Counterfactual Learning for Audio-Visual Segmentation Make-an-audio: Text-to-audio gen- eration with prompt-enhanced diffusion models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.380217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:25.010580Z digest=sha256:62affa9d27b7ee0bb2b5880622345bd85ba9ab6693bc0bad3a1c6fb95e4574b9

Observation 22163450-fa2b-4b2c-b5fa-5ccb7264a524 · outbound

This paper cites Discovering Sounding Objects by Audio Queries for Audio Visual Segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Discovering Sounding Objects by Audio Queries for Audio Visual Segmentation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:25.169098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:25.169098Z digest=sha256:e50fcf55e51f4ef0776ed68b5ec0cdce89e44179c8c4c8ec88cda8f847442517

Observation 68326339-9d84-4988-b9c3-9c53df2ff659 · outbound

This paper cites Adaptive mixtures of local experts.Neu- ral computation, 3(1):79–87, 1991.

Implicit Counterfactual Learning for Audio-Visual Segmentation Adaptive mixtures of local experts.Neu- ral computation, 3(1):79–87, 1991

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.371190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:25.330912Z digest=sha256:0ed2bed4d83e7081cba1f0722c5e59bed944c0c62592c014311ca5e70a6e2abe

Observation 6c4c9f0d-b966-4443-b3cb-2633c9986a55 · outbound

This paper cites Counterfactually augmented event matching for de-biased temporal sentence grounding.

Implicit Counterfactual Learning for Audio-Visual Segmentation Counterfactually augmented event matching for de-biased temporal sentence grounding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.362368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:25.559673Z digest=sha256:042cf2baae4a5c00eba6ac7bc41dccdd97cf7a887e2676fee1eb735500123d1c

Observation 993faee2-f328-4068-9126-e24260f76fb1 · outbound

This paper cites Learning to visually localize sound sources from mix- tures without prior source knowledge.

Implicit Counterfactual Learning for Audio-Visual Segmentation Learning to visually localize sound sources from mix- tures without prior source knowledge

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.353458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:25.728411Z digest=sha256:1e2c23662929dc7f79c0c1e32f656a1f664cf07005a61a1219fb1d5c2d083857

Observation dd417650-4dd9-4d51-b7d5-2064dcdfbd7e · outbound

This paper cites Segment Anything.

Implicit Counterfactual Learning for Audio-Visual Segmentation Segment Anything

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:25.802594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:25.802594Z digest=sha256:b45b0323df34c7ddb7e6d7aba7b18f80138c1a5be4544836b6de6b17b0a45994

Observation e63916f1-06da-44b1-9a0a-9fb875251c68 · outbound

This paper cites The singular value decompo- sition: Its computation and some applications.

Implicit Counterfactual Learning for Audio-Visual Segmentation The singular value decompo- sition: Its computation and some applications

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.344825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:25.806450Z digest=sha256:728f9a36aa5aa46810133c8e0028ef323f7eb345c7f8ebccfb2ba6ad5fab820a

Observation 474a0f74-f627-4521-b5f7-23d07f5141c8 · outbound

This paper cites Improving vision and language concepts understanding with multimodal counterfactual samples.

Implicit Counterfactual Learning for Audio-Visual Segmentation Improving vision and language concepts understanding with multimodal counterfactual samples

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.335417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:25.892511Z digest=sha256:f8019682f672bd974e71bde5e31daaf50b68eeeb1c8c1dfb860a8697536b0479

Observation 408135f8-3618-45b7-bf4c-cdddf4533d7d · outbound

This paper cites Selm: Selective mechanism based audio-visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Selm: Selective mechanism based audio-visual segmentation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.325244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:26.058249Z digest=sha256:292831a5c49b470a92206b74adb902866d19674daa24d09ccfd57025428145e2

Observation cb089aa7-6c78-41ad-8705-7df5c63882fe · outbound

This paper cites Catr: Combinatorial-dependence audio-queried transformer for audio-visual video segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Catr: Combinatorial-dependence audio-queried transformer for audio-visual video segmentation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.315333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:26.185331Z digest=sha256:f480f2fc39269451011ad655bc8fd7e96767abf1a6628aea64f05d09868ee994

Observation 4bda86bf-c11d-4172-9739-69e6e3a052d8 · outbound

This paper cites Dice Loss for Data-imbalanced NLP Tasks.

Implicit Counterfactual Learning for Audio-Visual Segmentation Dice Loss for Data-imbalanced NLP Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:26.309876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:26.309876Z digest=sha256:ad88c9c3176e2290e056c62d4bf1aed1a061332ed207063ef01a4df95a277881

Observation dc6ce626-8fa9-44c3-8a01-d4ab5140b521 · outbound

This paper cites Qdformer: Towards ro- bust audiovisual segmentation in complex environments with quantization-based semantic decomposition.

Implicit Counterfactual Learning for Audio-Visual Segmentation Qdformer: Towards ro- bust audiovisual segmentation in complex environments with quantization-based semantic decomposition

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.305293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:26.423838Z digest=sha256:218acc094f499a7bf78fdf3301a09b17de40ed6e15967ce812ca4466d7c619fb

Observation 07b5e024-bb4d-4dfc-8e01-9cb7c3ac55b4 · outbound

This paper cites Bavs: bootstrapping audio- visual segmentation by integrating foundation knowledge.

Implicit Counterfactual Learning for Audio-Visual Segmentation Bavs: bootstrapping audio- visual segmentation by integrating foundation knowledge

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:38.084022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:26.526751Z digest=sha256:5d36a952a2f7f9df929fe7970dce702dfdcb31f819fab7ffb718fba2f49faa5a

Observation cf977619-d618-4910-b7d3-7d9da6e76caf · outbound

This paper cites Benchmarking au- dio visual segmentation for long-untrimmed videos.

Implicit Counterfactual Learning for Audio-Visual Segmentation Benchmarking au- dio visual segmentation for long-untrimmed videos

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:37.893219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:26.586204Z digest=sha256:409d1e4a4c1ec4f64b15aa3eec8b25fed7714dd0dcc93aeaf6b83ddf84162d72

Observation b31a92b5-bc64-4781-84c0-0900693af95b · outbound

This paper cites Pay attention to mlps.

Implicit Counterfactual Learning for Audio-Visual Segmentation Pay attention to mlps

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:37.688252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:26.693841Z digest=sha256:4914348015d3b76df34179cd38bae991290eac7d7c488f123aa27d52efa54af1

Observation 83406f5f-7bb9-4ab9-8444-f8fed48f887d · outbound

This paper cites Audio-aware Query-enhanced Transformer for Audio-Visual Segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Audio-aware Query-enhanced Transformer for Audio-Visual Segmentation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:26.828928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:26.828928Z digest=sha256:9ff463133968f5f3df1403e805c977468e10b2d1df5940244d7c6349da78b0d0

Observation 2c8f30a5-6a29-4d72-bdfd-1edacc700458 · outbound

This paper cites Annotation-free Audio-Visual Segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Annotation-free Audio-Visual Segmentation

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:24:31.926622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:26.936793Z digest=sha256:255c1ce35537f3737822b0f807b42457194dc1daf1be8d424524bf6c1b1143fa

Observation 625ce05b-28ab-495a-b273-91ed366358e5 · outbound

This paper cites Audio-visual segmentation via unlabeled frame exploitation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Audio-visual segmentation via unlabeled frame exploitation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:37.518602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:27.047524Z digest=sha256:ade304191db5edfd2cc286e547125f27a21ae8300a56150ddb64445c455ab31b

Observation c7579be3-fa9c-49a3-be3a-10f786024ce9 · outbound

This paper cites Cross-modal causal relational reasoning for event-level visual question answer- ing.

Implicit Counterfactual Learning for Audio-Visual Segmentation Cross-modal causal relational reasoning for event-level visual question answer- ing

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:37.282649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:27.109504Z digest=sha256:e8fd2180cf03abaf5ff4f077ee91a97a457e238e7c9c2ca47303dad48c787def

Observation a06da19b-3e21-4ffb-b1d3-ac841e088740 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Implicit Counterfactual Learning for Audio-Visual Segmentation Swin transformer: Hierarchical vision transformer using shifted windows

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:27.197551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:27.197551Z digest=sha256:fc40790c3f6e6dc4604dbf3b4ed0a20a1d8ade0a2890562eae28b3d4ef4811b8

Observation 13a86cc8-f036-4e73-9a36-d75679152aaa · outbound

This paper cites Step- ping stones: a progressive training strategy for audio-visual semantic segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Step- ping stones: a progressive training strategy for audio-visual semantic segmentation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:37.131875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:27.320757Z digest=sha256:8d52607ddc5466577b5e2e0d2f3ce40a56a8f9fe367b1a370586cba5eb96ad08

Observation b64d31e0-fc4b-46aa-a4d2-719d938986bf · outbound

This paper cites T-vsl: Text-guided visual sound source localization in mixtures.

Implicit Counterfactual Learning for Audio-Visual Segmentation T-vsl: Text-guided visual sound source localization in mixtures

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.993803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:27.386920Z digest=sha256:f1726bbad4ad738942c91a804fe069da4dec4a568c52ff4aea9940606f331ea0

Observation 076f56ba-c832-4b60-b184-fb3d02046c3e · outbound

This paper cites Contrastive Conditional Latent Diffusion for Audio-visual Segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Contrastive Conditional Latent Diffusion for Audio-visual Segmentation

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:24:31.766124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:27.507201Z digest=sha256:efe548236bc0315056c8e1dee0ad7f53d2b27a76809d516117e8c88664bd5f78

Observation 055e248f-983a-45a3-8192-908b2a38d400 · outbound

This paper cites Multimodal variational auto-encoder based audio-visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Multimodal variational auto-encoder based audio-visual segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.818435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:27.623102Z digest=sha256:b0abb38cbc13ded78586aa4bc449901628158830b6cfee5b57dcbc7628061826

Observation 6bd68199-fb71-4bcd-95f9-262304aef98b · outbound

This paper cites Weakly-supervised audio- visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Weakly-supervised audio- visual segmentation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:27.707306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:27.707306Z digest=sha256:8e67fa6057ebff551d02876929d7a4b82c515794a7dfea4a31bdda000007a348

Observation 679f4081-6369-4b62-8236-66eb37155659 · outbound

This paper cites Audio-visual grouping net- work for sound localization from mixtures.

Implicit Counterfactual Learning for Audio-Visual Segmentation Audio-visual grouping net- work for sound localization from mixtures

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.647904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:27.831618Z digest=sha256:4f8d9f587012371141855022bbbaa59aecf7a5840304e8187208dbc6e179ca24

Observation b1a01a9e-85cd-4fce-8f44-da6ae9d201d1 · outbound

This paper cites The book of why: the new science of cause and effect.

Implicit Counterfactual Learning for Audio-Visual Segmentation The book of why: the new science of cause and effect

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.501190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:27.932310Z digest=sha256:fad1cdea5fbd71b1bde072b90643c32549766a8c52a424941d98af050fc93227

Observation bcabcb0c-4af6-402f-ac10-adf0ba432cfd · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Implicit Counterfactual Learning for Audio-Visual Segmentation Learning transferable visual models from natural language supervi- sion

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:28.047650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:28.047650Z digest=sha256:c6af273ecfa21c21421cd68686ea692af30b976bffced07bfbb2e22f30ecf9a4

Observation 550e563d-bb74-48bb-8da8-98f2af648330 · outbound

This paper cites Zero-shot text-to-image generation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Zero-shot text-to-image generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:28.155249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:28.155249Z digest=sha256:be712709c4b38f81ea7d3fb3927fae85e96582bb56029a3db6da2cd87e56b197

Observation a0617c11-9af7-4ca2-aed7-99a53b509b4b · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Implicit Counterfactual Learning for Audio-Visual Segmentation High-resolution image synthesis with latent diffusion models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.393079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:28.258287Z digest=sha256:3ca2d85f669807e6d95cf48c85373ae6c1d71222024100082b90d097b3ac49cd

Observation 5e9251e1-5e22-4c0c-87d5-841fd30864bb · outbound

This paper cites Focal loss for dense ob- ject detection.

Implicit Counterfactual Learning for Audio-Visual Segmentation Focal loss for dense ob- ject detection

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:28.331897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:28.331897Z digest=sha256:255b24de1e854b515559cce49db362b2e2e01e570eadc7bd32e69a6b2edd00d9

Observation af45f266-bf12-4f4b-8d4c-88b9a904e7be · outbound

This paper cites The wasserstein distance and approxi- mation theorems.

Implicit Counterfactual Learning for Audio-Visual Segmentation The wasserstein distance and approxi- mation theorems

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.249885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:28.439113Z digest=sha256:bab61c2d21f76e2b03b0ccfe7e896a9f41e75eeb036c7986888b2b19d06642e1

Observation f7ab43a7-705a-400e-92ea-0d19b46e369c · outbound

This paper cites Better aggregation in test-time augmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Better aggregation in test-time augmentation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:36.097862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:28.519677Z digest=sha256:955fe4b2938a356b93e952096ac3422df8d01fad694a1c6db805d4372d7c077a

Observation ffb19ff6-108c-4007-954d-feb8189ad7d3 · outbound

This paper cites A mathematical theory of commu- nication.

Implicit Counterfactual Learning for Audio-Visual Segmentation A mathematical theory of commu- nication

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.925906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:28.634597Z digest=sha256:de14e57c6e7422be18dfb506f40782dd3c6d00be440376853f1b1fcc43216750

Observation 99535a89-fd96-4c7c-bd56-9077b0df0bef · outbound

This paper cites Looking similar sounding different: Leveraging counterfactual cross-modal pairs for audiovisual representa- tion learning.

Implicit Counterfactual Learning for Audio-Visual Segmentation Looking similar sounding different: Leveraging counterfactual cross-modal pairs for audiovisual representa- tion learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.804483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:28.704904Z digest=sha256:b0c7141889d60ac003d7a7ba34e436346b8fdd72b3c1638a5db03cb486d23b4a

Observation 23110967-391c-4226-986c-db595b6da29e · outbound

This paper cites 3d audio-visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation 3d audio-visual segmentation

Reference 59

Resolution
verified exact
raw_fallback, observed 2026-08-06T13:24:31.614704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:28.777772Z digest=sha256:c96e2a6b0957deb1cc176e020e8c171c9d607774b6d9454640ec330d6a001bb8

Observation e2a02759-95dd-43e7-ac93-50bfdd89d177 · outbound

This paper cites Unveiling and mitigating bias in audio visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Unveiling and mitigating bias in audio visual segmentation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.702149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:28.858392Z digest=sha256:a49c9d072a8c4c9e73b95a31fd82e1b77d3a594f883ab8dab6ffb51eb4b3544a

Observation 15377690-589c-47c1-8b04-87f7639ab53a · outbound

This paper cites Learning audio-visual source localization via false negative aware contrastive learning.

Implicit Counterfactual Learning for Audio-Visual Segmentation Learning audio-visual source localization via false negative aware contrastive learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.572731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:28.979464Z digest=sha256:a9c0b2371002d38c7e0acf00176b4356b0b00bbb0d28092128de0e4758a5cbf0

Observation 2515ccf1-998f-4f4d-ac57-f3a328a0e817 · outbound

This paper cites Language-guided audio-visual source separation via trimodal consistency.

Implicit Counterfactual Learning for Audio-Visual Segmentation Language-guided audio-visual source separation via trimodal consistency

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.417838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:29.061286Z digest=sha256:3a577b619a4df496bb7375b7017ee1ba4876d450c2a66abb20f181fe362bec10

Observation 0c2e86cd-6ec9-4654-a9a2-c8867e76ca0a · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Implicit Counterfactual Learning for Audio-Visual Segmentation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:29.167740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:29.167740Z digest=sha256:20e97e0c87a1631680f94134fdecdfbad778c4f6f4b835413ce09642649640a5

Observation 8a2ca42f-3d2f-4d93-a5d3-fb2d02c7faac · outbound

This paper cites Neural discrete representation learning.

Implicit Counterfactual Learning for Audio-Visual Segmentation Neural discrete representation learning

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:35.232776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:29.269444Z digest=sha256:11151eeaf71d1000d2a8f562f9572fb15a2967761f40971f9624e168a5519d99

Observation f2efaf7a-148a-4976-ac2b-eb2b10eaeb84 · outbound

This paper cites Counterfactual cycle-consistent learn- ing for instruction following and generation in vision- language navigation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Counterfactual cycle-consistent learn- ing for instruction following and generation in vision- language navigation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.985658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:29.340354Z digest=sha256:61222f4051f3a4f547301fd4a26fdc612ba08d5c474bd08d21a52573dc0177a8

Observation dfe75355-433d-4e55-8086-5c7bb61ea5b4 · outbound

This paper cites Vision-and-language naviga- tion via causal learning.

Implicit Counterfactual Learning for Audio-Visual Segmentation Vision-and-language naviga- tion via causal learning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.817968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:29.450485Z digest=sha256:5b0606f13675deae194bb6f34da8bf351394f65611ea11b25e958d5e68b84aab

Observation 6b0363dd-d240-4a07-b11c-99257ee012d2 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Implicit Counterfactual Learning for Audio-Visual Segmentation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:29.575827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:29.575827Z digest=sha256:4672d3b96155bc9cb5ba770b4a353873a3968626b54d20d952f90dae670a3148

Observation 57326a31-7fed-419a-9e03-f04f782c88ca · outbound

This paper cites Pvt v2: Improved baselines with pyramid vision transformer.

Implicit Counterfactual Learning for Audio-Visual Segmentation Pvt v2: Improved baselines with pyramid vision transformer

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.591109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:29.651671Z digest=sha256:8bdfd6f5fe721f89d79d2473e22e8dca5b85dcf8a204d472395b6b79ab2bb07d

Observation 2d8e8a11-5f77-4b10-8d4c-a21c080864b3 · outbound

This paper cites Drivedreamer: Towards real-world- drive world models for autonomous driving.

Implicit Counterfactual Learning for Audio-Visual Segmentation Drivedreamer: Towards real-world- drive world models for autonomous driving

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:29.738949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:29.738949Z digest=sha256:5578f7a2a2b1919024b04eff48767004ed99a68c34515d33d5feab2dd59d94f3

Observation b00d8a91-7e04-44c4-be8e-414d4a62e7f1 · outbound

This paper cites Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer.

Implicit Counterfactual Learning for Audio-Visual Segmentation Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:24:31.372411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:29.839727Z digest=sha256:d90f70ba0516a00aa874143941386129170b254cd472e28f601dd2acb86739dd

Observation d84413af-577e-40c0-bc13-1a6e731270b3 · outbound

This paper cites Prompting segmentation with sound is gen- eralizable audio-visual source localizer.

Implicit Counterfactual Learning for Audio-Visual Segmentation Prompting segmentation with sound is gen- eralizable audio-visual source localizer

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.366851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:29.911767Z digest=sha256:b232d432a50bcaec61cb8604c9bd9ab2dffdeb2dda1ec6bb88f6694d7047b4cf

Observation 2e1162cf-9d17-4e68-b0cc-8e8d71258c23 · outbound

This paper cites Can textual semantics mitigate sounding object segmentation preference? In European Conference on Com- puter Vision, pages 340–356.

Implicit Counterfactual Learning for Audio-Visual Segmentation Can textual semantics mitigate sounding object segmentation preference? In European Conference on Com- puter Vision, pages 340–356

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.206530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:30.002229Z digest=sha256:0ba5740e81f550a51a46d8284d0aecde4bb26b4be58e90af0098594b912fbeb3

Observation 145afdec-8d79-44d4-a233-1db959720e1c · outbound

This paper cites Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:34.033283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:30.078314Z digest=sha256:1a868441680c06948b86526953d6468d6854af53c6ee702a4e2f309ef0f4aa25

Observation 2f93e96f-c19c-41b4-8489-c070a09d9865 · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

Implicit Counterfactual Learning for Audio-Visual Segmentation VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:30.193705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:30.193705Z digest=sha256:70da2cf0f87433fce341599466e0438e4ef29f7bb3a2013d8f701ecfbfc314b1

Observation 25e4326e-a60c-4e3f-b7ee-99d4bb643f4a · outbound

This paper cites Referred by multi-modality: A unified tem- poral transformer for video object segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Referred by multi-modality: A unified tem- poral transformer for video object segmentation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.919685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:30.264676Z digest=sha256:4eccdec02f6419abc714f58f08f1e40100d7a7b6442b106a6e283e878412b03a

Observation 5434c471-de08-4491-804c-0091f009c7c1 · outbound

This paper cites Cooperation does matter: Exploring multi-order bilateral relations for audio- visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Cooperation does matter: Exploring multi-order bilateral relations for audio- visual segmentation

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.757091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:30.395792Z digest=sha256:876f9ba260e2bfba296fb2a71cccc6fa7c3e257c9509864a6418abf817094459

Observation c1f36da8-5497-4e1f-9186-b2268aa953a2 · outbound

This paper cites Revisiting counterfactual prob- lems in referring expression comprehension.

Implicit Counterfactual Learning for Audio-Visual Segmentation Revisiting counterfactual prob- lems in referring expression comprehension

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.617230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:30.482824Z digest=sha256:bfbd6232595623c717a01a675ce869177b5745d1baf184a247ef8f70015f298f

Observation fc20ea18-d27a-4797-a966-e1b8bd72dcee · outbound

This paper cites Losh: Long-short text joint prediction network for referring video object segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Losh: Long-short text joint prediction network for referring video object segmentation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.459830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:30.591093Z digest=sha256:09ea38d2508fa702c7d274cc98f9bf4683508361a8248d15b4b11121c09b1eb3

Observation 9c325e08-f56f-4645-b594-f7a34b8e575a · outbound

This paper cites Discovering the real association: Multimodal causal rea- soning in video question answering.

Implicit Counterfactual Learning for Audio-Visual Segmentation Discovering the real association: Multimodal causal rea- soning in video question answering

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.321674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:30.648645Z digest=sha256:0dd5416c1141b5a2078e10221afe26ebce0bda04c4b8a1d049f4b89bde879536

Observation de25b491-8657-4d1f-8b76-893d79779aad · outbound

This paper cites Weakly- supervised mirror detection via scribble annotations.

Implicit Counterfactual Learning for Audio-Visual Segmentation Weakly- supervised mirror detection via scribble annotations

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:33.147692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:30.711366Z digest=sha256:afb6ebc9599b79f5469168793cffbc0c1cc5c4fd42b46077ace5c8e321e7a59e

Observation 7410b16a-d8ab-4b01-b633-cbee76be50da · outbound

This paper cites Heterogeneous experts and hierarchical perception for un- derwater salient object detection.

Implicit Counterfactual Learning for Audio-Visual Segmentation Heterogeneous experts and hierarchical perception for un- derwater salient object detection

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:32.969158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:30.787086Z digest=sha256:455f9cdf234a5b98e1e7057daa82483924dcfb631361bf3a0cb191e86f73d146

Observation 5310ff46-d21b-4fc5-836b-00592f7203dc · outbound

This paper cites Causal intervention for weakly- supervised semantic segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Causal intervention for weakly- supervised semantic segmentation

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:32.874213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:30.860023Z digest=sha256:6db4606dc4265bd68ee8d4cb9e69b943bdcb39da234d6f6dd33270bcb70e24b1

Observation 919094cc-387d-44d6-9e58-d36771852dd4 · outbound

This paper cites Audio–visual segmentation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Audio–visual segmentation

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T13:24:30.918766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:24:30.918766Z digest=sha256:ec7c6606dd2658f17b35e62b881f33bb56eb01633faab281e87c963a7ea6b099

Observation ed7a1761-fd70-4dab-bcb5-a3e28308dedc · outbound

This paper cites Audio-visual segmentation with semantics.

Implicit Counterfactual Learning for Audio-Visual Segmentation Audio-visual segmentation with semantics

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:32.591154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:31.109915Z digest=sha256:6f5571c51c81c8a30eb6dc8b8d82aa60f5099489f9a31fe7ef50a0afbf2502e5

Observation cc711852-876c-4aff-83cc-ffd8923a2ae3 · outbound

This paper cites Exploring pre-trained text- to-video diffusion models for referring video object segmen- tation.

Implicit Counterfactual Learning for Audio-Visual Segmentation Exploring pre-trained text- to-video diffusion models for referring video object segmen- tation

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:32.424693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:31.180804Z digest=sha256:bee6148d47092e2b145a90d33a7ba345650eb56f2f99f24fd5416dbf975c343e

Observation 9abe49c0-d912-40e2-9f47-04af5b5e656a · outbound

This paper cites 2, 5, 6, 8.

Implicit Counterfactual Learning for Audio-Visual Segmentation 2, 5, 6, 8

Reference 403

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:24:32.709159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T13:24:31.029697Z digest=sha256:57be4c5f1b8f324e64856750830183318ef0ec9a9160e0f76eb39855b80570f6

Pith citing papers

No inbound Pith citation observations are available.