Pith. sign in

Paper Citation Record · LEDGER

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation

As of 11 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2501.13667.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.13667 v5

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:47:59.927188Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy38
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4e1c3f6c-bbc6-4ba5-8b13-9314fc777a8e · outbound

This paper cites Xmem++: Production-level video segmentation from few annotated frames.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Xmem++: Production-level video segmentation from few annotated frames

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:01.036615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.638768Z digest=sha256:9e25ffb5d87c0209b280cf4964ee2ffaefcd48b88734ad000bdac45c13eff452

Observation bbc84f67-177f-4a4d-9f71-db45046f149d · outbound

This paper cites RefVOS: A Closer Look at Referring Expressions for Video Object Segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation RefVOS: A Closer Look at Referring Expressions for Video Object Segmentation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.644137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.644137Z digest=sha256:7f892a1a22db8799d70eae782db9bc52bf671775bd22dd4a74b897eacfb323f6

Observation 6bf26444-f89a-4366-b158-ec586a8a771a · outbound

This paper cites End-to-end referring video object segmentation with multi- modal transformers.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation End-to-end referring video object segmentation with multi- modal transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:01.019502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.649698Z digest=sha256:37234241516a03e7796cb6dcc363bb67c3787f3a3226a4a32482745aa0749a08

Observation fdc8ee81-920a-4426-9562-f90f532b5172 · outbound

This paper cites End-to-end referring video object segmentation with multi- modal transformers.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation End-to-end referring video object segmentation with multi- modal transformers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:01.001502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.654829Z digest=sha256:0cd414daee84bb1b833cf8a7307d61dfbdd55a00b9e8ba487e5be576e7174857

Observation 987ce447-7a12-4674-bc10-519771e68145 · outbound

This paper cites End- to-end object detection with transformers.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation End- to-end object detection with transformers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.984032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.659696Z digest=sha256:1eb76f8f9e26d7058a71f6a2ae79af24237a2efe6a4cf355bcfce57ae72efdc9

Observation 755fb4b8-5a47-4de1-9e55-e8b1463f784c · outbound

This paper cites Xmem: Long- term video object segmentation with an atkinson-shiffrin memory model.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Xmem: Long- term video object segmentation with an atkinson-shiffrin memory model

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.966930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.664751Z digest=sha256:ff1c62768fd320772daff5d9da375288db80bf89b3a4dbbc2d1eaa5be2d599b2

Observation ff279dd1-d879-4b85-adb0-7c78be7d3270 · outbound

This paper cites Putting the object back into video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Putting the object back into video object segmentation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.950457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.670130Z digest=sha256:2de884b793855cf13cb658adb11a433add03b71b41c70b2c5cbbe8c46a358441

Observation 4546e337-6ea6-4a03-8f0c-ec6d3e3a81d0 · outbound

This paper cites Segment and Track Anything.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Segment and Track Anything

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.674573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.674573Z digest=sha256:28481cae3ec9869469429f93c985268eb17e4b9374aedf7eee40f6cf38ddc643

Observation 055a0632-966b-491b-95e1-cc2a2fbb4493 · outbound

This paper cites Unsupervised Cross-lingual Representation Learning at Scale.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Unsupervised Cross-lingual Representation Learning at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.679455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.679455Z digest=sha256:715ab286487a9b59b19dc6a900d4cec771af70475d325d55f41885d0630c4786

Observation ff018f14-f355-42b4-b98e-04a998bfec9a · outbound

This paper cites Vision-language transformer and query generation for refer- ring segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Vision-language transformer and query generation for refer- ring segmentation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.932309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.684899Z digest=sha256:74d77d81453a6cd94e4809465c0c6cc86d8c02c085834849c110de47a3c98cbc

Observation 0fa86d86-2c75-4729-a592-455d0adf2496 · outbound

This paper cites Mevis: A large-scale benchmark for video segmentation with motion expressions.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Mevis: A large-scale benchmark for video segmentation with motion expressions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.913517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.690233Z digest=sha256:ad746891a0ce1ff5213f88a219c07d546fd4af5fae3002efeaf6efd4b17e22ca

Observation 5971daea-c441-47f0-b780-85b500c13e4d · outbound

This paper cites Language-bridged spatial-temporal interaction for referring video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Language-bridged spatial-temporal interaction for referring video object segmentation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.896976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.695176Z digest=sha256:130bb11a57153e9afbf424decbe922cebddc7e9ef725ea44757ef099613f7bce

Observation f94ac849-ef94-41d6-ab2c-b0ef10281cb4 · outbound

This paper cites Unified embedding alignment for open-vocabulary video instance segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Unified embedding alignment for open-vocabulary video instance segmentation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.881133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.700026Z digest=sha256:41d08f9b8da60bcbbb05581feeec735914e33a15e77bd314017d2b3025d35898

Observation 64500324-77b2-49a1-b524-4c5921b9300e · outbound

This paper cites Html: Hybrid temporal-scale mul- timodal learning framework for referring video object seg- mentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Html: Hybrid temporal-scale mul- timodal learning framework for referring video object seg- mentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.865317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.704741Z digest=sha256:373726f30cbd958ce890314a7b8cc10f3d9762d936e0286f913b29456b41aa0e

Observation 160ea41c-b019-473c-9647-601d3cdd12c5 · outbound

This paper cites Decoupling static and hier- archical motion perception for referring video segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Decoupling static and hier- archical motion perception for referring video segmentation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.848532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.709388Z digest=sha256:66d01d6c88c49c8f2bc8f88e2794cd9dbace9cae0a9a834ded2f9eda50fe3e93

Observation 7161f8c8-2472-4e51-8763-388d1fb61b66 · outbound

This paper cites Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object Segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object Segmentation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.714033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.714033Z digest=sha256:37308340138ea81a2769021165f5f00d6e5bb1fb06e44581ca8abb5ee8f76df8

Observation 69c6c784-e421-4ba8-806c-7309be4bb745 · outbound

This paper cites Segment anything in high qual- ity.Advances in Neural Information Processing Systems, 36,.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Segment anything in high qual- ity.Advances in Neural Information Processing Systems, 36,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.718988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.718988Z digest=sha256:698f2afab9892d408b55c2f144878827a17ab589f8601cbefbacb4bd36882980

Observation dfe47324-1f8c-4d7a-aaad-9288711f0ff0 · outbound

This paper cites Video object segmentation with language referring expressions.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Video object segmentation with language referring expressions

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.815303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.723933Z digest=sha256:8fae4071a3954cc25104814001867b6a61f393c43eb29f2b78a53c2de53747c6

Observation 4e3d5cfb-376c-4f68-a950-94455ae29c78 · outbound

This paper cites Segment any- thing.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Segment any- thing

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.797935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.728435Z digest=sha256:a32a42c3dc493a26f4028423d1c2f029cdc8d8a1a80dd51323787a6182724be5

Observation b89bf872-4ce4-4000-89c9-2bd0ba6b7222 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Lisa: Reasoning segmentation via large language model

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.780736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.733316Z digest=sha256:bdb2d692f6bf02fb821d532412396b4a686e25744e691678c415aaed67525000

Observation ced4a7ee-c08d-48ca-af59-a115b9104a63 · outbound

This paper cites Learning to learn better for video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Learning to learn better for video object segmentation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.764725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.737900Z digest=sha256:9c25cdf797701b626fb5fd4b6785523ca62b1e7c0ab05588954521fad2f062ab

Observation 328add4e-aec4-408f-9b64-772b9545be10 · outbound

This paper cites an unresolved cited work.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-10T15:48:00.749283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.742330Z digest=sha256:46b735ada7b7a6097a2707d7c5cb4b0e8ea880cf50f54fbc57edb91a93f100ab

Observation 5b791fd6-2d71-4e4c-a4a1-6985029cb2a2 · outbound

This paper cites Bidirectional correlation-driven inter-frame inter- action transformer for referring video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Bidirectional correlation-driven inter-frame inter- action transformer for referring video object segmentation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.733566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.747025Z digest=sha256:bbd28c0d3e274c7dde76f9a61f688c8e12b7ea06926a4e3dbb5758f463ec11c0

Observation 9701dfa4-8e8e-4eef-886e-cd73a0ea8d1c · outbound

This paper cites You only infer once: Cross-modal meta-transfer for referring video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation You only infer once: Cross-modal meta-transfer for referring video object segmentation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.717969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.751536Z digest=sha256:2ee8a51bd6d62a678babc75657cc58d690902db2b91f669b7c845bd91085d0a3

Observation e5c6324e-ed1f-44bc-9d23-12fb89929b46 · outbound

This paper cites RefSAM: Efficiently Adapting Segmenting Anything Model for Referring Video Object Segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation RefSAM: Efficiently Adapting Segmenting Anything Model for Referring Video Object Segmentation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.756198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.756198Z digest=sha256:e016022594483e24472a82f2f1495cbb20aba05c4b44535be024f561370526f0

Observation 3e9f7077-2c38-4e6e-99f5-159afcd2d098 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.761037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.761037Z digest=sha256:31532c7c3b63c1610e7fae2e09a6d3882b65ce6c1408d32e6eff41658974d3d9

Observation fd5bc5ec-19e1-40f6-b456-e6f50d648946 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.765578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.765578Z digest=sha256:2e620a319c412bd4e03be9d448e00f769aa28ca547135c2c13d2bad44a9a848e

Observation 230a054c-9723-4a4a-93ea-8bf4aac4c325 · outbound

This paper cites Decoupled weight decay regularization.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Decoupled weight decay regularization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.770551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.770551Z digest=sha256:cb7b8c248aea17c6a6dcb230c96ee1d35521654fad43308c6d6641d53f4dfbae

Observation d947f84c-2012-4c4a-94d4-3e539f1415bd · outbound

This paper cites Soc: Semantic-assisted object cluster for referring video object segmentation.Advances in Neural Information Processing Systems, 36, 2024.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Soc: Semantic-assisted object cluster for referring video object segmentation.Advances in Neural Information Processing Systems, 36, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.682289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.775051Z digest=sha256:1f22348a214fced4804c40e68faa6a0ae0b095382d6ee0cd93d0101e6c6da417

Observation 50a38455-8280-4a0f-9131-321f32e3234f · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Generation and comprehension of unambiguous object descriptions

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.779445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.779445Z digest=sha256:83d896ced3108c116b60babfe3fc70106a935dc2b5d953f3bac1e0b34ba9c917

Observation bd1b8c88-a717-449d-a81c-135e74c8e4ab · outbound

This paper cites Visual-textual capsule routing for text-based video segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Visual-textual capsule routing for text-based video segmentation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.655677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.784047Z digest=sha256:08248759282434ccb968d77ddace436de45e01c4c43d58dba871fb8b2542cc01

Observation 38e25771-d08d-407a-a24a-19c3b8c25675 · outbound

This paper cites Spectrum-guided multi-granularity referring video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Spectrum-guided multi-granularity referring video object segmentation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.640795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.788681Z digest=sha256:f83987940672be791d77fac7685f86a96b8ccecce0ff68d0eecd037165a85cfb

Observation 04df44cf-137c-43d2-9023-03e4e5a2394f · outbound

This paper cites V-net: Fully convolutional neural networks for volumetric medical image segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation V-net: Fully convolutional neural networks for volumetric medical image segmentation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.624776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.793025Z digest=sha256:619230e71ecbaf92e8585f2e7743a7a96a093b34e644c6d387d11b7f2be2a306

Observation ad7d8cdc-6449-4416-bddc-15badaeee255 · outbound

This paper cites Video object segmentation using space-time memory networks.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Video object segmentation using space-time memory networks

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.609152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.797817Z digest=sha256:305ddb1d2d3a9f8992225ef2619fdb78cab75f3e46f6989da4149207b1406293

Observation b476880e-13e2-4bc3-ba23-53d3216bf272 · outbound

This paper cites Semantic and sequential alignment for referring video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Semantic and sequential alignment for referring video object segmentation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.593023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.802324Z digest=sha256:9674ac74ecd85b3b486630548cc84803c6e7646b144d8c0e1863e23b9929a1df

Observation a2d4b4e7-3df5-46c6-88d9-7eb0fb05c1c2 · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation The 2017 DAVIS Challenge on Video Object Segmentation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.807087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.807087Z digest=sha256:132347d6f06a85ace1f5a803cfcd01886684c90e995445b6f8e4f9c884157cff

Observation e384604e-d98d-4289-b000-5cd9177d1b1e · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Glamm: Pixel grounding large multimodal model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.811931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.811931Z digest=sha256:e134ff649e5c7f868ea8dd2c7b21b7f5634675df17526761affd301549e251c7

Observation 36430e9f-d503-4473-b3e4-ed7f044aac35 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation SAM 2: Segment Anything in Images and Videos

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.817076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.817076Z digest=sha256:edeb88ca7185b322e235fa367a5d931d184d3f3a183331f490db3999c88cdfe2

Observation 1dd84681-af88-483f-aac6-521cc1ddbc2e · outbound

This paper cites Cus- tomized sam 2 for referring remote sensing image segmenta- tion.arXiv preprint arXiv:2503.07266, 2025.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Cus- tomized sam 2 for referring remote sensing image segmenta- tion.arXiv preprint arXiv:2503.07266, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.821941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.821941Z digest=sha256:dc7300a1fc14403ce8052951a82695cadd70ec86f3682489708446382cdb86b8

Observation 92ed7359-2cfb-4b8d-985e-b3e434a333e0 · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Urvos: Unified referring video object segmentation network with a large-scale benchmark

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.567693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.826708Z digest=sha256:72a8b73e24c8575ec76aa7b278814e005bdc8b9d5feb9cb02275c4d2ce259e9f

Observation 94e8d278-d03b-4049-8de5-aca91dbf4e0b · outbound

This paper cites Temporal collection and distribution for referring video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Temporal collection and distribution for referring video object segmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.552878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.831198Z digest=sha256:d6aefe895683d4bf2eeee8b282c7f7487c3c6aeb0a57fb40a3eabfd51a888dab

Observation 3065cf5b-7afd-4526-833f-13bb09643fd2 · outbound

This paper cites Samrs: Scaling-up re- mote sensing segmentation dataset with segment anything model.Advances in Neural Information Processing Systems, 36, 2024.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Samrs: Scaling-up re- mote sensing segmentation dataset with segment anything model.Advances in Neural Information Processing Systems, 36, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.537906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.835790Z digest=sha256:a9774338a61f277f5fb21150a1576338e17999032cd069b9f6df9d89c370ab59

Observation a1d6c23c-9577-466b-b38a-327dcc6a4f30 · outbound

This paper cites Asymmetric cross-guided attention network for actor and ac- tion video segmentation from natural language query.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Asymmetric cross-guided attention network for actor and ac- tion video segmentation from natural language query

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.521970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.840270Z digest=sha256:639dad524cd05943e5c23a18937709d0e4d25b89bec3dc08a1e2472cf757cbbd

Observation 5959649d-ce2c-46b4-9e05-e8c3ca604f26 · outbound

This paper cites Image as a foreign language: Beit pretraining for vision and vision- language tasks.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Image as a foreign language: Beit pretraining for vision and vision- language tasks

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.504760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.845067Z digest=sha256:d521ec1f8f912ae4f2e8037f838563dda7979662458bd2bd0202ebd923745bea

Observation 8f1e15ea-a785-4cc4-ace7-b3b1e43c0986 · outbound

This paper cites HyperSeg: Towards Universal Visual Segmentation with Large Language Model.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation HyperSeg: Towards Universal Visual Segmentation with Large Language Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.849646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.849646Z digest=sha256:1f6dfb168d8929989f7686b8e0628db354303e58903e7230e71ff42c0d9d6c72

Observation 99a2b93d-69c9-4020-8819-a18dc653a03a · outbound

This paper cites Multi-level representation learning with semantic alignment for referring video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Multi-level representation learning with semantic alignment for referring video object segmentation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.488432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.854413Z digest=sha256:2f5386f723ba13b372e1989fcccd992d292cf882f49f77ec1724361c717f01e9

Observation dcaacb23-6094-40d8-a4eb-12d4d8299c9d · outbound

This paper cites Onlinerefer: A simple online baseline for referring video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Onlinerefer: A simple online baseline for referring video object segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.471795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.858650Z digest=sha256:1f06b54ce3897dc5b2e67ad1a154397f65b893013c01defdda2e86b4d486818a

Observation 27dad8b8-5237-45d9-b191-642f246c9a2d · outbound

This paper cites Language as queries for referring video object segmen- tation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Language as queries for referring video object segmen- tation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.455585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.863235Z digest=sha256:989bd158d7c981a509fe0d70aca826469b7b7dcbca56341c6746571a64ad1504

Observation af69f5b1-cba9-4ea5-bb47-dfc980093eef · outbound

This paper cites Logiczsl: Exploring logic- induced representation for compositional zero-shot learning.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Logiczsl: Exploring logic- induced representation for compositional zero-shot learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.438701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.868031Z digest=sha256:efcb0e791ccae2f013c2d27a2ac912540351f27cce8ae6a32890e546ca696221

Observation 3690eaf6-802f-4159-9288-f20a2e6b9847 · outbound

This paper cites Efficientsam: Leveraged masked image pretraining for efficient segment anything.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Efficientsam: Leveraged masked image pretraining for efficient segment anything

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.422489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.873004Z digest=sha256:fc29832b99244fbd9b493e0ef508440be74c111f386be50c99bb1b099264786a

Observation d52ffb01-c86d-4286-9866-779efd545e76 · outbound

This paper cites u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.877607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.877607Z digest=sha256:7dc7a28d8685f8dfaf898a1f49e64e1bdc2ce9eaeb9e49a0ee4763ccc3bb5d62

Observation de3c20d8-e016-42fd-8746-f16c4ab22481 · outbound

This paper cites Visa: Reasoning video object segmentation via large language models.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Visa: Reasoning video object segmentation via large language models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.405801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.882837Z digest=sha256:5448e1b016672cb389a5181e4b268d97dc551f6225ed7d24b25606b70ee1e82d

Observation f36f6979-76b4-4c7c-9717-e82879d29075 · outbound

This paper cites Referred by multi-modality: A unified tem- poral transformer for video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Referred by multi-modality: A unified tem- poral transformer for video object segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.390008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.887488Z digest=sha256:bc0575857df128ffc56e599995f8c68a53cf8c93b3477c118f38d1a6509e9e61

Observation 39fab523-0cbf-4cde-a3ea-496bc9ffb894 · outbound

This paper cites Modeling context in referring expres- sions.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Modeling context in referring expres- sions

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.372685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.892095Z digest=sha256:ea04a2ebb891f30c2080655f2ce03a6a68f82a44fa734e571dd2638932d3e0ab

Observation cc420157-1065-4e3b-a348-dada5e697f3f · outbound

This paper cites A Simple Baseline with Single-encoder for Referring Image Segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation A Simple Baseline with Single-encoder for Referring Image Segmentation

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-10T15:48:00.007696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.897025Z digest=sha256:517b73c1830637422246317d1f64ec4e233d8d296cce98d2cc31005678debc4c

Observation 959aed02-6d41-47d3-9e66-7763d8836a05 · outbound

This paper cites Losh: Long-short text joint prediction network for referring video object segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Losh: Long-short text joint prediction network for referring video object segmentation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.901850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.901850Z digest=sha256:99ad2fecd011cbbf85ed9c22440d9553d28dbdeaa3b99ea7e4abad768e7e613d

Observation 015c654e-2242-41cf-a9d4-161540a8f6cb · outbound

This paper cites Surgicalsam: Efficient class prompt- able surgical instrument segmentation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Surgicalsam: Efficient class prompt- able surgical instrument segmentation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.906540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.906540Z digest=sha256:4350eeebcb07b05caa7a3ca8c035b75e224d9d2d60f9621aed9cc0972ba8c0fe

Observation 42fa24c1-7262-43ff-885e-fd50ec4fb259 · outbound

This paper cites Faster Segment Anything: Towards Lightweight SAM for Mobile Applications.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Faster Segment Anything: Towards Lightweight SAM for Mobile Applications

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.911548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.911548Z digest=sha256:11220cc7409a10f612c0b209771606d12694bde991a69e6c079ca65684806310

Observation 747b87e0-fc36-405d-a4e0-52a21ae4e60d · outbound

This paper cites EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.916575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.916575Z digest=sha256:166805236b9611fc0728cc022c64b35fec6a6338b2bbd17e0176cce899cdd347

Observation a569ecdb-9f60-4755-9229-2bedc8ca6b6c · outbound

This paper cites Deformable detr: Deformable transformers for end-to-end object detection.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Deformable detr: Deformable transformers for end-to-end object detection

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T15:47:59.922407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:47:59.922407Z digest=sha256:92114758bc24999ebf24161408439a74318085d2880e6235aa8874796663c20a

Observation 0cd8a6d0-c9f7-4798-b2df-3d797451fb83 · outbound

This paper cites Exploring pre-trained text- to-video diffusion models for referring video object segmen- tation.

MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation Exploring pre-trained text- to-video diffusion models for referring video object segmen- tation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:48:00.326524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:47:59.927188Z digest=sha256:48c4e5a82ea2561f094a5e58f98e5996cd342ff0f859305eb06f1ab9a371f214

Pith citing papers

No inbound Pith citation observations are available.