Pith. sign in

Paper Citation Record · LEDGER

InterRVOS: Interaction-aware Referring Video Object Segmentation

As of 8 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 1 inbound Pith citation observation for arXiv:2506.02356.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02356 v3

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:29:29.724751Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T01:50:54.242508Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.213452Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dee763a0-162e-4fdf-803b-272cdbdf7042 · outbound

This paper cites One token to seg them all: Language instructed reasoning segmentation in videos.

InterRVOS: Interaction-aware Referring Video Object Segmentation One token to seg them all: Language instructed reasoning segmentation in videos

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:34.253558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:29:27.623640Z digest=sha256:80c12d81433fab84e0cb9cb542cf2599a27561637401bb0bae2fb135dd26a791

Observation 85184742-8ac3-4641-a574-aecbcd50b8c5 · outbound

This paper cites End-to-end referring video object segmentation with multimodal transformers.

InterRVOS: Interaction-aware Referring Video Object Segmentation End-to-end referring video object segmentation with multimodal transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:34.057605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:29:27.674228Z digest=sha256:b4bafd7d4b2fed3a11615dcca0bf04eac41c344f0d4edfc87ba57c789d3211ef

Observation 5a381820-99cb-45bc-9192-783b00a298f8 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

InterRVOS: Interaction-aware Referring Video Object Segmentation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:27.777419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:27.777419Z digest=sha256:dfd0bb66b693a791a29ee6ce82e9c17e2b128148df854f58136809e5febd2253

Observation 9368ec74-88b7-4f2f-a2c8-0c690f01fbda · outbound

This paper cites Vision-language transformer and query generation for referring segmentation.

InterRVOS: Interaction-aware Referring Video Object Segmentation Vision-language transformer and query generation for referring segmentation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:33.862292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:29:27.834762Z digest=sha256:5110d0524b1ace9a4729ae4caff300f5cc6e42487d8dabbef1af988db77f23fe

Observation baae64bd-df07-4a70-b212-5abaec036c7b · outbound

This paper cites Mevis: A large-scale benchmark for video segmentation with motion expressions.

InterRVOS: Interaction-aware Referring Video Object Segmentation Mevis: A large-scale benchmark for video segmentation with motion expressions

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:33.640172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:29:27.917308Z digest=sha256:0b1df32a94e361625fd7af20315c812cf62b4526d679f67b8e000561c3935f0d

Observation b91dbc1a-cbb9-4d1f-aa8b-a8e0ac7e1dd8 · outbound

This paper cites Moma: A multi-object multi-action dataset for understanding human activities.

InterRVOS: Interaction-aware Referring Video Object Segmentation Moma: A multi-object multi-action dataset for understanding human activities

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:33.410861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:29:27.989653Z digest=sha256:f86455d5f103db0327e52f2be3320336ee1ea95dcc43255d3c373ff27487abef

Observation bf4a6e33-31c5-4bfb-ad76-b6f4018096bb · outbound

This paper cites Actor and action video segmentation from a sentence.

InterRVOS: Interaction-aware Referring Video Object Segmentation Actor and action video segmentation from a sentence

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:33.306269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.076797Z digest=sha256:1486582657e935d5ad603c4a4e6e38e98cd73ee2d918cd2aa415047e616d5535

Observation 52848177-7ddd-4880-9b09-00a850dcaec7 · outbound

This paper cites The llama 3 herd of models.

InterRVOS: Interaction-aware Referring Video Object Segmentation The llama 3 herd of models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:33.101988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.133567Z digest=sha256:9a41ab9ce011d5222cda8a6f63ef33f0ec6d11832ccafb980ea68e938e837791

Observation f51f17a5-0d00-409d-b20a-94e1e8dfbb97 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

InterRVOS: Interaction-aware Referring Video Object Segmentation Lora: Low-rank adaptation of large language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:28.190322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:28.190322Z digest=sha256:7daf9b52812d31598c1ffa3a53fd84f65ccba1f9b5d2ea4fd4ea8aade5798f55

Observation 82a39cc5-de55-4213-afa2-c3f0483ebaad · outbound

This paper cites GPT-4o System Card.

InterRVOS: Interaction-aware Referring Video Object Segmentation GPT-4o System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:28.254083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:28.254083Z digest=sha256:78362abd7e795eb68f92e150fcb258da4689921074c1e1dcf1c089422e417cd0

Observation 5ad1d67c-d91c-4f6d-a940-9f0cc036ade9 · outbound

This paper cites Action genome: Actions as compositions of spatiotemporal scene graphs.

InterRVOS: Interaction-aware Referring Video Object Segmentation Action genome: Actions as compositions of spatiotemporal scene graphs

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:32.908278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.343745Z digest=sha256:cedcdcc5706d59e31ff460f58ef51e8b57d359a43a14ba7a25098bbaa195fc4b

Observation 86063e0b-4cee-4535-af7a-e12a2afdbae7 · outbound

This paper cites Video object segmentation with language referring expressions.

InterRVOS: Interaction-aware Referring Video Object Segmentation Video object segmentation with language referring expressions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:32.619250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.412310Z digest=sha256:f8d2cec383458ae6824b3a3f9b74f1937720933c82b0a6c3b959c7e67ea5bdef

Observation 3c42c50f-28d6-4f2a-a0c1-5d2179bed05c · outbound

This paper cites VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos.

InterRVOS: Interaction-aware Referring Video Object Segmentation VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:28.478864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:28.478864Z digest=sha256:94566592dbfbd68734c62a6624a2bec07e91e013793f19980a0bf77d5512adbc

Observation 9dd12239-b4e5-43e3-95f1-eab0bf2524d0 · outbound

This paper cites Visual instruction tuning.

InterRVOS: Interaction-aware Referring Video Object Segmentation Visual instruction tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:28.562199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:28.562199Z digest=sha256:44ff65fd7cafa4c49609d929e93f3fd57ee2eefc1e3f7f7b3f8852a324c9bf0e

Observation b8e90b91-2bcd-4c5a-b288-d8694652df00 · outbound

This paper cites Spectrum-guided multi-granularity referring video object segmentation.

InterRVOS: Interaction-aware Referring Video Object Segmentation Spectrum-guided multi-granularity referring video object segmentation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:32.416144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.655247Z digest=sha256:83395ce351eeec15e501ebb296fc154930b907c0df9eda8128fc9171ed9f6ba1

Observation 945be114-8625-4e57-9bbe-9fc57529e925 · outbound

This paper cites Refer-youtube-vos: A dataset for video object segmentation with language referring expressions.

InterRVOS: Interaction-aware Referring Video Object Segmentation Refer-youtube-vos: A dataset for video object segmentation with language referring expressions

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:32.178119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.718185Z digest=sha256:fbc674fecbc4734a26dd61d634cf2a30cbe9a702dfb6c889bf8525720f131a7b

Observation 80d7bdab-9f61-4f66-a496-b15b3a3355e5 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

InterRVOS: Interaction-aware Referring Video Object Segmentation SAM 2: Segment Anything in Images and Videos

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:28.763180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:28.763180Z digest=sha256:527efffeb4d4c46ab69b13b80467c677e7c6e0ab9b688eec59fbc053b8c64072

Observation db77b5ba-1be2-4a7a-9791-269ea4c563db · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark.

InterRVOS: Interaction-aware Referring Video Object Segmentation Urvos: Unified referring video object segmentation network with a large-scale benchmark

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:31.861439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.841080Z digest=sha256:8be1c718c561f6d74f01cd76d4d7e3b0bffe7efa47bb9431d9011b10e8987321

Observation 2aa89ca2-753c-40bd-b9e7-699f0531a945 · outbound

This paper cites Video relationship detection.

InterRVOS: Interaction-aware Referring Video Object Segmentation Video relationship detection

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:31.474955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.914632Z digest=sha256:0747dbb594c55290372bf9b9eb4e32dfc58de42cabe8a2c1fe71335f0e098f37

Observation 71050332-79e4-42da-8df2-789af37553ba · outbound

This paper cites Annotating objects and relations in user-generated videos.

InterRVOS: Interaction-aware Referring Video Object Segmentation Annotating objects and relations in user-generated videos

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:31.094806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.981295Z digest=sha256:4518b6a978a6ab996c90f9ebdec079d95a490ed874ced278763b96af6e3d9149

Observation e30af10b-5310-4e50-a02a-6535d18e2ab7 · outbound

This paper cites Yfcc100m: The new data in multimedia research.

InterRVOS: Interaction-aware Referring Video Object Segmentation Yfcc100m: The new data in multimedia research

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:30.865552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:29:29.039012Z digest=sha256:e845b47cacb05b18f941531f6c45bb3d752eb5c7e0b97566524ae546438884fb

Observation 5cb9c047-76ec-408a-8c66-0f6c3378604a · outbound

This paper cites SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation.

InterRVOS: Interaction-aware Referring Video Object Segmentation SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.102921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.102921Z digest=sha256:d69de8137bf90b33d36f411f556e6baf60865e322a2a124b842e13ed95df79a1

Observation 51c3e5c4-7511-4120-a3b2-65b765c8c9b2 · outbound

This paper cites ViLLa: Video Reasoning Segmentation with Large Language Model.

InterRVOS: Interaction-aware Referring Video Object Segmentation ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.149297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.149297Z digest=sha256:0719a8c21ecafb3641d9bfe87bcbf17f275662a3a5544d0376c331176581bf54

Observation e701675b-9556-4572-b972-89f1ac135800 · outbound

This paper cites Language as queries for referring video object segmentation.

InterRVOS: Interaction-aware Referring Video Object Segmentation Language as queries for referring video object segmentation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:30.644265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:29:29.205824Z digest=sha256:10ace87b1c054049a453886b971c637a07884c858e2c47f8787a7e9699c97758

Observation 759cfd64-7b93-4d6e-9e2e-1b9fa7556250 · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

InterRVOS: Interaction-aware Referring Video Object Segmentation STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.288871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.288871Z digest=sha256:c5c6fb4749a21593b55432590affadfef0de4290f6e44ea2308883542445f6e0

Observation 3efc5d46-5ba5-4ef7-8314-11c0529b260c · outbound

This paper cites Visa: Reasoning video object segmentation via large language models.

InterRVOS: Interaction-aware Referring Video Object Segmentation Visa: Reasoning video object segmentation via large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:30.477138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:29:29.323770Z digest=sha256:b394e000f9eb086bc0a372e01dfb733c5d2262772674afa2760664f073772fa5

Observation 651c05f3-96ee-4b49-9b2e-fae4fe859f67 · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

InterRVOS: Interaction-aware Referring Video Object Segmentation Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.406573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.406573Z digest=sha256:3c910a79cc9e7e03e23468c3fe11dbb71160069fe1118d5d9089b658d8606ff5

Observation f2776946-a890-4711-adbb-6403f95f64e9 · outbound

This paper cites Decoupling Static and Hierarchical Motion Perception for Referring Video Segmentation.

InterRVOS: Interaction-aware Referring Video Object Segmentation Decoupling Static and Hierarchical Motion Perception for Referring Video Segmentation

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:30.142164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T11:29:29.474932Z digest=sha256:bff9b36c114fcb0744cc2ed9f03d2825ea347939bca4f383910b12c2ef80093e

Observation 9034104e-4cb4-4125-b42f-18ec53fd65aa · outbound

This paper cites write newline.

InterRVOS: Interaction-aware Referring Video Object Segmentation write newline

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.534945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.534945Z digest=sha256:a1ab4cb0703e5848e04c3d0c434ec9236def6d12c14d8991cc68f4c9f3715750

Observation 1d30c372-044f-4464-b690-037358a9bc37 · outbound

This paper cites @esa (Ref.

InterRVOS: Interaction-aware Referring Video Object Segmentation @esa (Ref

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.594298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.594298Z digest=sha256:b4666ef46d4c323781404fbe5424632039daff025b76d26411be0c4d834448f2

Observation c13ea854-4ba4-4081-8f51-25305ab4c2b2 · outbound

This paper cites an unresolved cited work.

InterRVOS: Interaction-aware Referring Video Object Segmentation Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.642842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.642842Z digest=sha256:ac808c53ead722f7498e2ab040e21b2a8cc7ecbaae32d28a7c1cfe527f628203

Observation ce1d1cc0-77ce-4bad-b412-6a3faf66d064 · outbound

This paper cites A child helping another child with a backpack.

InterRVOS: Interaction-aware Referring Video Object Segmentation A child helping another child with a backpack

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.724751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.724751Z digest=sha256:63db45a2b72a1fe3055ec4e1d27ad425c81346ca17597c87682593f1b7ddeb13

Pith citing papers

Observation 66f5d8ba-b7bd-4887-b54e-b3e839842af9 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models InterRVOS: Interaction-aware Referring Video Object Segmentation

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.214944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:846544d7afd40b3568cbdabb5dda24e670bdd9c57444352e61ff58db10287970