Pith. sign in

Paper Citation Record · LEDGER

X2SAM: Any Segmentation in Images and Videos

As of 6 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2605.00891.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.00891 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-09T20:47:06.698475Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact13
  • verified fuzzy54
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 189a533b-936a-4955-8d15-8da0e87348e6 · outbound

This paper cites Qwen Technical Report.

X2SAM: Any Segmentation in Images and Videos Qwen Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:01:05.310910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:597a75bed4726c4ab4a9333f6503cfabcdcb8c809556bd8363ad994405f514fa

Observation ee46ec15-f3da-40f0-beb6-71859a1d303e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

X2SAM: Any Segmentation in Images and Videos LLaMA: Open and Efficient Foundation Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:01:05.240526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:8bfa1a27c6f7a39cdcad776f8053f0033ceb51460cf49ec7ed09a44d4b33382f

Observation 112e7e39-cdcb-4a74-bed1-3d0328ae86cd · outbound

This paper cites Learning transferable visual models from natural language supervision.

X2SAM: Any Segmentation in Images and Videos Learning transferable visual models from natural language supervision

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.536243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:a670a6ada194ad85d9d1505e5b469b00adda29a50a3d7a895df344e4f3c45212

Observation d661328e-2fc3-4b90-9f66-18185220cd93 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

X2SAM: Any Segmentation in Images and Videos Scaling up visual and vision-language representation learning with noisy text supervision

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.498229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:3957fcbfab46b93e26cd5e348e2662973672690ed79a53a0773277cbf3ee8d84

Observation 1a9c4b47-0ff5-4ef6-9da4-ce5691db9d97 · outbound

This paper cites Show, attend and tell: Neural image caption generation with visual attention.

X2SAM: Any Segmentation in Images and Videos Show, attend and tell: Neural image caption generation with visual attention

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.539604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:8fb1a4284ed7651fce9d1bc3f6b7c24920308892168c543ec2cc4a157f087e4b

Observation efbb2c68-ac56-46c8-9834-5c011cd288bd · outbound

This paper cites Vqa: Visual question answering.

X2SAM: Any Segmentation in Images and Videos Vqa: Visual question answering

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.494715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:5f0a1c26a0ab7de16c400e2aae1bee1e4df2dc037df649a3a410f288de417ec5

Observation 8a666af3-d6f6-40ee-95c4-935b067aeaff · outbound

This paper cites Language-based image editing with recurrent attentive models.

X2SAM: Any Segmentation in Images and Videos Language-based image editing with recurrent attentive models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.501296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:b10291bc2eb10b7b15834ad75c6a431d5bb62b12ac5e584a552bffee39d4898d

Observation 73d5b524-410c-4d1c-a816-5c329c2a46cb · outbound

This paper cites Segment anything.

X2SAM: Any Segmentation in Images and Videos Segment anything

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.504364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:ab296f3cd5b3693640eaccc9d9893edb215711e47d3354a77d11a4551f53341e

Observation 57511c89-3be4-4704-b86d-a1b3c87186d1 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

X2SAM: Any Segmentation in Images and Videos SAM 2: Segment Anything in Images and Videos

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:01:05.356359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:d0148bf9a53330ce422a847af92310285e9a63688af66e707337a565f06e7b64

Observation 9a1001aa-4c7b-42f2-8239-3d622cc1ab50 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

X2SAM: Any Segmentation in Images and Videos Lisa: Reasoning segmentation via large language model

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.510978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:09cb30d01dee4bfcc59bbd8f6e949686fed3e0d8a1601c46b0f7de3d7901a9d5

Observation 83a0d30e-093f-4cc7-8039-969fd5a5b955 · outbound

This paper cites Visa: Reasoning video object segmentation via large language models.ECCV.

X2SAM: Any Segmentation in Images and Videos Visa: Reasoning video object segmentation via large language models.ECCV

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.476693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:10c71fd9ca0637bf8d735f358c366bc967c55bf0df2715b857c0326c4dc01f85

Observation d02ec514-18dc-494d-abc9-69eba5c3511d · outbound

This paper cites One token to seg them all: Language instructed reasoning segmentation in videos.NeurIPS.

X2SAM: Any Segmentation in Images and Videos One token to seg them all: Language instructed reasoning segmentation in videos.NeurIPS

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.480003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:0f51759127669c7f4ac5b2a47e93716c5d592edc535223af5a938b5c89af178e

Observation ef49ba74-329f-4077-92a8-5222ed275bfb · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

X2SAM: Any Segmentation in Images and Videos Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.483298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:5a9b9619f05408522f047cb501405480d5c6e7917ef2ec1f6d1abbcbe7d5ec1b

Observation 843c713f-225c-47ee-b7f7-0f8ab3275bbc · outbound

This paper cites Improved baselines with visual instruction tuning.

X2SAM: Any Segmentation in Images and Videos Improved baselines with visual instruction tuning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.514451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:c28ba831d23f8ce5bbc5bb30746b6f32d1773578aa56b6138b1c59de14a3cf04

Observation c95311b6-83b1-40e2-8254-c9c2a49a22b4 · outbound

This paper cites Tarvis: A unified approach for target-based video segmentation.

X2SAM: Any Segmentation in Images and Videos Tarvis: A unified approach for target-based video segmentation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.521963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:55df0c84e274b561baa033daf2977fac446ed9aafc2c741c69dfd49b0a1dd9d6

Observation aeba86b3-ffcb-4e68-b3b7-56a9ca1126c1 · outbound

This paper cites Oneformer: One transformer to rule universal image segmentation.

X2SAM: Any Segmentation in Images and Videos Oneformer: One transformer to rule universal image segmentation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.528451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:42713ebe7057bc849f4cc1f543664981c3cf7eb7cfdc8cee73dc97a4425dbe11

Observation 2df7d21c-1b2b-4b3b-9f7c-1ade3bd75d29 · outbound

This paper cites Omg-seg: Is one model good enough for all segmentation? InCVPR, pages 27948–27959.

X2SAM: Any Segmentation in Images and Videos Omg-seg: Is one model good enough for all segmentation? InCVPR, pages 27948–27959

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.682610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:6f0b9b142510c6536780fa8b4f5a0171dd0cdc49a311c61cf0ab08442ac7abe3

Observation c7c79173-c57b-4342-bb2f-8493a602bb72 · outbound

This paper cites Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding.NeurIPS, 37:71737–71767.

X2SAM: Any Segmentation in Images and Videos Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding.NeurIPS, 37:71737–71767

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.678888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:0be7835c3f575d80e65472c00a894aeace4168bf72aa4b67ee1a1ac1b0b41dc0

Observation 00024a58-bc61-40ad-8391-b81693f6e39e · outbound

This paper cites Temporal memory attention for video semantic segmentation.

X2SAM: Any Segmentation in Images and Videos Temporal memory attention for video semantic segmentation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.686154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:519a2008f7bcda49edddccbfdfbfe57ae67376d7634bc92ea5b0c7ecfadd7fe8

Observation f512585c-0589-43c6-ae74-50191e885586 · outbound

This paper cites Video k-net: A simple, strong, and unified baseline for video segmentation.

X2SAM: Any Segmentation in Images and Videos Video k-net: A simple, strong, and unified baseline for video segmentation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.693255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:f0d2fdcf9106723aadf0a538467f585a628564b24ff66f2a27651123daaece46

Observation af656a47-b134-4013-b0d3-5e5c5ed71abe · outbound

This paper cites X-sam: From segment anything to any segmentation.

X2SAM: Any Segmentation in Images and Videos X-sam: From segment anything to any segmentation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.667366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:bb17b9bfa8942e02d0413e205ebb15a8d27012040a9b26dd6ab4b7c3503a7d97

Observation 6974a5d1-5310-4dd2-baeb-57ed19f590ea · outbound

This paper cites Visual instruction tuning.NeurIPS, 36:34892–34916.

X2SAM: Any Segmentation in Images and Videos Visual instruction tuning.NeurIPS, 36:34892–34916

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.671074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:821c2bfe477802b02e7d596c8736f8acdb1bcefa600ade6abe5761303a369803

Observation 64660bb0-4f1c-4366-924a-f0b000bbab47 · outbound

This paper cites Llavanext: Improved reasoning, ocr, and world knowledge.

X2SAM: Any Segmentation in Images and Videos Llavanext: Improved reasoning, ocr, and world knowledge

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.518515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:9faf2346781311d5591011b6c2377fc02cd97eec1eb98dd6e0c1ee4a733e224f

Observation 02ce4c79-32d9-49c6-941c-5a1e0530e1d5 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

X2SAM: Any Segmentation in Images and Videos LLaVA-OneVision: Easy Visual Task Transfer

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:01:05.327718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:c55734b69c96d860de0bf576a8ab5ee63cef0fd96bcd66430b1ebebc14578c79

Observation c1ff8307-044d-4a7e-be60-6a2c5bfcf9cf · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101.

X2SAM: Any Segmentation in Images and Videos How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.660214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:f909b33370416379f7089aeb477c2b8e3f19d863d00e1990cad48c42c29e7675

Observation 896b17eb-d5e7-4d8f-939c-172841dd903a · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

X2SAM: Any Segmentation in Images and Videos Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:01:05.456688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:5df895560af8ab00676b7148609c9db5e8194d4871d5897dea19ccf857df5394

Observation 8a919a35-e3c3-4889-8b0a-33b4989eb4b0 · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

X2SAM: Any Segmentation in Images and Videos Glamm: Pixel grounding large multimodal model

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.652775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:6350ec94cfe5f8a25e3dde0f1b3a4ca56881447a5685b23cac502a0e208d550d

Observation 0196f2a1-bdb6-4125-8449-43c5cd76b280 · outbound

This paper cites Videoglamm: A large multimodal model for pixel-level visual grounding in videos.CVPR.

X2SAM: Any Segmentation in Images and Videos Videoglamm: A large multimodal model for pixel-level visual grounding in videos.CVPR

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.656287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:3a3769d175d26f422c31edfe4b9b38475dcc36de5973f7fcc837329e4638de54

Observation af0980ee-7629-4667-9e72-759526988977 · outbound

This paper cites Psalm: Pixelwise segmentation with large multi-modal model.

X2SAM: Any Segmentation in Images and Videos Psalm: Pixelwise segmentation with large multi-modal model

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.663736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:619b48ad012c8ccdd31b4b01add24c4ee4ea77705d00f68b1329ea4f03f89179

Observation bba4b529-0457-41d8-84ff-7f7bc8d4facf · outbound

This paper cites HyperSeg: Towards Universal Visual Segmentation with Large Language Model.

X2SAM: Any Segmentation in Images and Videos HyperSeg: Towards Universal Visual Segmentation with Large Language Model

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:01:05.436176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:d294cd8c4e526c88fa320ac37cda8b5a704013da824fab21f6310ffe8a1aabaf

Observation b81759f9-1147-4d4a-819b-121ec2ec0a0b · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

X2SAM: Any Segmentation in Images and Videos Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:39:22.805579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:198703fdf4865ade7c301af2ef36370eadd07a47e5e35a44bf44f1fb42776f3e

Observation f2b0c637-ef50-418c-992d-c7bd6da00cb1 · outbound

This paper cites Qwen3-VL Technical Report.

X2SAM: Any Segmentation in Images and Videos Qwen3-VL Technical Report

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:01:05.383742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:0e8b1afec102a99e7b2299610445533b4168c5c4744a8f9935a0b372f54a1600

Observation ec346914-1640-4eef-afb5-371ccabdbf0e · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

X2SAM: Any Segmentation in Images and Videos Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:29:05.422888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:3071878cc69d7abf4d84334de2e9b3f63e1bce71bc5af2bfddc7f8180e3c94c4

Observation 426e0b0d-f9b4-4e99-9fbf-3abad408de99 · outbound

This paper cites V-net: Fully convolutional neural networks for volumetric medical image segmentation.

X2SAM: Any Segmentation in Images and Videos V-net: Fully convolutional neural networks for volumetric medical image segmentation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.674761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:d0d32c843b11bb5a1adbb9c7a728d184295a8905408b064845f0672d6a839e66

Observation 0e9620a1-d871-43b7-bbab-e92b044516c7 · outbound

This paper cites Improving language understanding by generative pre-training.

X2SAM: Any Segmentation in Images and Videos Improving language understanding by generative pre-training

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.635492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:25cda972aac71fd8eada0535a5aca43272864be1d1c8abaad92af12e844b6c50

Observation 48ae41de-27a5-4684-b2ce-c0d10650fb2b · outbound

This paper cites Focal loss for dense object detection.

X2SAM: Any Segmentation in Images and Videos Focal loss for dense object detection

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.639720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:0713cd35fcdc770c2167d20be52e07e3155d48f4ff1ba8c984d738be78bd2153

Observation 5714d776-dbe0-48c9-a429-e69de5135a60 · outbound

This paper cites Large-scale video panoptic segmentation in the wild: A benchmark.

X2SAM: Any Segmentation in Images and Videos Large-scale video panoptic segmentation in the wild: A benchmark

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.627025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:7450bfff7495587b07b7fe3419d69f5c23aaa35f4266773e6c0496889de189ee

Observation ecd7e4e6-cd0b-47d2-be74-ce892fff8d39 · outbound

This paper cites Vspw: A large-scale dataset for video scene parsing in the wild.

X2SAM: Any Segmentation in Images and Videos Vspw: A large-scale dataset for video scene parsing in the wild

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.619200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:01335074e5bb659a87727c092ccdb3cfb6a263cafe1f0d2ab09313804c0f125b

Observation b59f63e8-6e66-4f52-ba88-c9b6cacba9d5 · outbound

This paper cites Video instance segmentation.

X2SAM: Any Segmentation in Images and Videos Video instance segmentation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.622865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:215458320bc20c326f90779ad4296f7cbac44e081cf6a434ebb793433c121e3f

Observation 9cf3bdca-9cc3-4380-b93c-f4fa1c233029 · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark.

X2SAM: Any Segmentation in Images and Videos Urvos: Unified referring video object segmentation network with a large-scale benchmark

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.631094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:e8fe2bf34951f178017d1cc9ab46f0644befc15e71cb232f8de89504b2be909b

Observation 2cb0fd27-26fd-4a0c-a948-00e046bcd75b · outbound

This paper cites YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark.

X2SAM: Any Segmentation in Images and Videos YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:01:05.428048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:7a47c64585ac8893b6f6421d4c54086e21cd9af27dd9d656abff420579432c19

Observation 1b3a79e1-793e-47b5-bfbb-8acd67dd255f · outbound

This paper cites A benchmark dataset and evaluation methodology for video object segmentation.

X2SAM: Any Segmentation in Images and Videos A benchmark dataset and evaluation methodology for video object segmentation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.642997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:b56a1be912d14713018e84509a6581de31b0c185adc3f6584de44942515e6e6f

Observation 979af393-7e8b-4cc5-8da4-0f631655a1df · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

X2SAM: Any Segmentation in Images and Videos Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.646842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:88e64fb28a16b9e6c9dd9a10baf88cca64f21167105f6ad88b3b56c08a6a6ed1

Observation 4ff49862-0d0c-4d00-bc18-6d5cc73c35f3 · outbound

This paper cites Gres: Generalized referring expression segmentation.

X2SAM: Any Segmentation in Images and Videos Gres: Generalized referring expression segmentation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.689595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:c7cdd7a6257c33a13b9988f72a49928d4ad1b9a23616df2f81c30a5c5da86546

Observation c65917f6-a72f-48a0-9a66-37df14f5e1ea · outbound

This paper cites Semantic understanding of scenes through the ade20k dataset.IJCV, 127(3):302–321.

X2SAM: Any Segmentation in Images and Videos Semantic understanding of scenes through the ade20k dataset.IJCV, 127(3):302–321

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.696795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:9d2f0890508be09a08614bebc6bc805d3dcce8e4b99aab2d17fa38013d60f5d1

Observation a8702369-aeaf-4897-814e-cf6494515560 · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

X2SAM: Any Segmentation in Images and Videos LoRA: Low-rank adaptation of large language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.598446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:7d3b3324b90f36364176180b09baa4f89d6887e3d031e4e6a0331056c67a97ba

Observation d1bf84a3-8b50-4e07-bb31-8e73d1a805c3 · outbound

This paper cites Decoupled Weight Decay Regularization.

X2SAM: Any Segmentation in Images and Videos Decoupled Weight Decay Regularization

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:01:05.228126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:1d40e1f04b329c1ef7cbb43259f49c33b32d99701c02cef19cfabc4985e1e20b

Observation 185942e1-e13f-4040-8fc0-7373101d478f · outbound

This paper cites Open-vocabulary panoptic segmentation with text-to-image diffusion models.

X2SAM: Any Segmentation in Images and Videos Open-vocabulary panoptic segmentation with text-to-image diffusion models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.580826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:0407a4e1de6ddd3228603e7be94fa07ee3032fd330c5dd00a176bb283414088b

Observation 5775c423-951a-4a6d-866a-f7911a22f73e · outbound

This paper cites Uniref++: Segment every reference object in spatial and temporal spaces.ICCV.

X2SAM: Any Segmentation in Images and Videos Uniref++: Segment every reference object in spatial and temporal spaces.ICCV

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.486963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:3a7b41e3530831102bcbf0533a4f06913e863127db103a8db7df64dbcd46465f

Observation ddd436b2-739a-471c-bb30-79590767642d · outbound

This paper cites Unipixel: Unified object referring and segmentation for pixel-level visual reasoning.

X2SAM: Any Segmentation in Images and Videos Unipixel: Unified object referring and segmentation for pixel-level visual reasoning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.491118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:d4e294c53ec79085bc1e1cc6f783982492ae4727753e959144da9bc3a5eb3f06

Observation 5442f6fb-7a08-4af4-be2b-a5640366d55e · outbound

This paper cites Segment everything everywhere all at once.NeurIPS, 36:19769–19782.

X2SAM: Any Segmentation in Images and Videos Segment everything everywhere all at once.NeurIPS, 36:19769–19782

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.507529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:1b1aa259dfb520e06decc5dd37b89d05df38cfd0ea881862bee83cab55a0a313

Observation 9bc49f70-74a6-4281-ae04-8ba5b8871e34 · outbound

This paper cites Language as queries for referring video object segmentation.

X2SAM: Any Segmentation in Images and Videos Language as queries for referring video object segmentation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.525512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:ea0d768eb74450abaa2f064e5893bd2e9f9c5c8bcdcbddb08abceb8da1a9253f

Observation 2af7add8-b3e8-4b47-9326-6b86450a8b53 · outbound

This paper cites LaSagnA: Language-based Segmentation Assistant for Complex Queries.

X2SAM: Any Segmentation in Images and Videos LaSagnA: Language-based Segmentation Assistant for Complex Queries

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:01:05.418347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:096d7d5bce2fd52ae0d6dab36c2facc5d965e068943e92d1a2e028d14f07f864

Observation fc734836-8251-4641-914b-d10e09e118b0 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InECCV, pages 216–233.

X2SAM: Any Segmentation in Images and Videos Mmbench: Is your multi-modal model an all-around player? InECCV, pages 216–233

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.532469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:a7d06850d03d918876e75d7bf17cc938bd764792c17783393aad0c6ee816e42c

Observation f03881bb-577e-44d7-8196-ef4ba1862f66 · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models.

X2SAM: Any Segmentation in Images and Videos Seed-bench: Benchmarking multimodal large language models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.584571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:f6a4d2f8fc4e62f224a894e384ce95e957dceb27b265f52df1d06fc1050b05ef

Observation 5fdedad0-1ec3-4043-96e2-c902833b1ce8 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

X2SAM: Any Segmentation in Images and Videos Evaluating Object Hallucination in Large Vision-Language Models

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:01:05.291358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:3b53240ac122e20c0cc1368860d8bab69acf364843fe7339e2ded336f9b24d0f

Observation 9068ae36-3707-43b3-ac8a-39d266f118ab · outbound

This paper cites A diagram is worth a dozen images.

X2SAM: Any Segmentation in Images and Videos A diagram is worth a dozen images

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.590763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:47dde56befe3305e3abaafb3d1872024c41e2d8bfb8ef7107dc94d007dd72386

Observation d7dc8b0b-ed25-4f49-90b6-46dbdaae1da6 · outbound

This paper cites Microsoft coco: Common objects in context.

X2SAM: Any Segmentation in Images and Videos Microsoft coco: Common objects in context

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.574063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:9422da21f6633a2c3370b4386da9186a0ccc70fb20086abddffef5994a8772ed

Observation 01815b34-207c-46b7-a5d7-16668649f050 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

X2SAM: Any Segmentation in Images and Videos Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.566380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:06e4bff793634106922b7483a4158d5f7f3d0ddd435faa926b6134d190ce02bd

Observation 671e29d7-b81b-4d47-b747-e2a39e9f0744 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

X2SAM: Any Segmentation in Images and Videos Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.570596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:5f25041eac084ef455834141cddd45a146867fc43b812bb8d42dfd6e456fd3a6

Observation 9f13f2ab-6e5e-450e-92a2-663be295ff92 · outbound

This paper cites Mlvu: Benchmarking multi-task long video understanding.

X2SAM: Any Segmentation in Images and Videos Mlvu: Benchmarking multi-task long video understanding

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.577708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:53bf5bd43bbd32c7f85e61fec35401258d76402115c6ea119f08d4f90afb8e2b

Observation 85d37fd3-c47c-45a1-8d4b-29d6ce4dc423 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.NeurIPS, 37:28828–28857.

X2SAM: Any Segmentation in Images and Videos Longvideobench: A benchmark for long-context interleaved video-language understanding.NeurIPS, 37:28828–28857

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.594663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:bddb98b4c28e4d31ff2536b9e4e3080d7f22ccc079092bcff0c45df81972255f

Observation 03bffbe7-5c54-49a6-8635-8d9a376f551f · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

X2SAM: Any Segmentation in Images and Videos Masked-attention mask transformer for universal image segmentation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.608794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:2cfd7c1ec7105c609a0fd5c2087b0efa5621218199f4ba0edcdc8ad30a0b7ed9

Observation 0d3a54b7-25d4-4926-8063-01edbde880b2 · outbound

This paper cites Cris: Clip-driven referring image segmentation.

X2SAM: Any Segmentation in Images and Videos Cris: Clip-driven referring image segmentation

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.562547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:7b5f5e3431b14119aa6f20a2ec1bd2abd2969ff3d77ebbe8b1c482c4a7214a1f

Observation 94abc434-0dfc-4bc3-ace5-c915712e418f · outbound

This paper cites PG-Video-LLaVA: Pixel Grounding Large Video-Language Models.

X2SAM: Any Segmentation in Images and Videos PG-Video-LLaVA: Pixel Grounding Large Video-Language Models

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:01:05.448353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:ce1de37b4008a709ad0df7690e762ebd3d4267b6f3be0fc0ba1864f4c7872427

Observation a2142cf6-256c-4bea-b154-a585cc137bd9 · outbound

This paper cites Videochat: Chat-centric video understanding.Science China Information Sciences, 68(10):200102.

X2SAM: Any Segmentation in Images and Videos Videochat: Chat-centric video understanding.Science China Information Sciences, 68(10):200102

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.552357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:54211cf6afaa6dd466aaa642a7acf39a02eddfec3cc4d3d0a586db972afdef6b

Observation 7feb229b-18d6-4698-89c7-3d85fe25ea84 · outbound

This paper cites Chat-univi: Unified visual representation empowers large language models with image and video understanding.

X2SAM: Any Segmentation in Images and Videos Chat-univi: Unified visual representation empowers large language models with image and video understanding

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.543460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:9c670a9fbe07a2dfb8bb8fec78e3761225c919bc1046c70e23a1ab7a4451bf31

Observation 167756bb-13c3-4fc1-b365-064c3411a38c · outbound

This paper cites Vila: On pre-training for visual language models.

X2SAM: Any Segmentation in Images and Videos Vila: On pre-training for visual language models

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T08:29:12.547625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T20:47:06.698475Z digest=sha256:96341006a05b6d4ac8bb4ecd714839ce095418f78a92dacac05af447ec1c707e

Pith citing papers

No inbound Pith citation observations are available.