Pith. sign in

Paper Citation Record · LEDGER

SAM 3: Segment Anything with Concepts

As of 5 August 2026, this Paper Citation Record lists 100 of 168 outbound references and 100 inbound Pith citation observations for arXiv:2511.16719.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.16719 v2

Coverage vector

measured 100 of 168 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T20:22:46.220021Z

measured 200 of 200 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 100 of 480 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T21:06:03.793460Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 168 outbound references displayed

  • verified exact32
  • verified fuzzy62
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

7
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation f7eaaba7-80d5-40fa-a4e9-01409a57e85c · outbound

This paper cites write newline.

SAM 3: Segment Anything with Concepts write newline

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.542845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:015aced6bb3a62ce3e11057ec00026319c36a9fc3584479ebe3162efeb30513b

Observation 0fd3a05d-19e2-40f6-bfaf-7bf887d25995 · outbound

This paper cites Greenhouse gas equivalencies calculator.

SAM 3: Segment Anything with Concepts Greenhouse gas equivalencies calculator

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.539950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:73238e10d51719dc222b544762346965dd924218dc549a9815a6bfc29ee865e4

Observation bc05edb3-0754-4fca-be81-836b309940bf · outbound

This paper cites Multi-label cluster discrimination for visual representation learning.

SAM 3: Segment Anything with Concepts Multi-label cluster discrimination for visual representation learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.261769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:c3de00a1a6df8f862530efc763be7fe69100658fd14d73b95a7fd1d9737dcf4b

Observation f7709340-2aa5-486c-9c89-e19951adb736 · outbound

This paper cites Burst: A benchmark for unifying object recognition, segmentation and tracking in video.

SAM 3: Segment Anything with Concepts Burst: A benchmark for unifying object recognition, segmentation and tracking in video

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.315170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:a615eda53464eada7ea3fe9cb0bbb6ef23e7eafd0563f0f233dbfce0aaa15d5f

Observation 4c8b77f4-a836-4cd7-b6d2-c92cf9ded9a3 · outbound

This paper cites Gmot-40: A benchmark for generic multiple object tracking.

SAM 3: Segment Anything with Concepts Gmot-40: A benchmark for generic multiple object tracking

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.309425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:8ecc70bee9c6490a2acb5294fd1d783f9c002f064f7a560aec7dd1b3d25610a3

Observation 2a4cb442-38fc-4d48-ae34-39900a057a93 · outbound

This paper cites DeepSea MOT: A benchmark dataset for multi-object tracking on deep-sea video.

SAM 3: Segment Anything with Concepts DeepSea MOT: A benchmark dataset for multi-object tracking on deep-sea video

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.218899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:b971fd2fb84889cf87611be8e2dae64b0bba86b7e9dc9b687e8b4c2268101eb3

Observation dd2d058e-26a5-4deb-ac77-25956a0821df · outbound

This paper cites Tracking without bells and whistles.

SAM 3: Segment Anything with Concepts Tracking without bells and whistles

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.320140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:51d5b7f96ece55bb27733cdf9171e0a2678306727eb2c7ff86ff1989b00905a2

Observation a5229b02-6507-4bc9-92e9-a3734517d11d · outbound

This paper cites Simple online and realtime tracking.

SAM 3: Segment Anything with Concepts Simple online and realtime tracking

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.343909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:23ab6cd9bd74c4e90b1b96b15b8bb40a034b46bf7150793c498814df4e105a97

Observation 024d1d5e-f874-45fa-8e01-a2f7a1bf7662 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

SAM 3: Segment Anything with Concepts PaliGemma: A versatile 3B VLM for transfer

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:25:11.400127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:6a12219cdd77ec806fa2258aa9c69625f47ad8ee735ab3a6c7ad4af0091b7278

Observation 84dfcf42-aef0-41a5-8fe4-7c0887d10576 · outbound

This paper cites YOLOv4: Optimal Speed and Accuracy of Object Detection.

SAM 3: Segment Anything with Concepts YOLOv4: Optimal Speed and Accuracy of Object Detection

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:25:11.417758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:517d8ba67472b5e014b0113d396f4a120ba1d707dead8f429fb2d3b9a9f62fd1

Observation 050aa7e8-efa4-4c69-9730-0ee1b4ba80d8 · outbound

This paper cites Window attention is bugged: How not to interpolate position embeddings.

SAM 3: Segment Anything with Concepts Window attention is bugged: How not to interpolate position embeddings

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.346435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:d5507e2d580d2e8bb1096cf5daa3dc3e48aaa47279fa22bdad5f41140678225c

Observation 865c7c00-f5dc-43fd-ad58-056324471fea · outbound

This paper cites Perception Encoder: The best visual embeddings are not at the output of the network.

SAM 3: Segment Anything with Concepts Perception Encoder: The best visual embeddings are not at the output of the network

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:25:11.582160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:428b1955dbc82c1155153f2d488734f07147b73032b10a7494993486a121ea60

Observation 3984b8c9-bfd6-4b9b-aba5-751c86b575a0 · outbound

This paper cites Align-detr: Enhancing end-to-end object detection with aligned loss.

SAM 3: Segment Anything with Concepts Align-detr: Enhancing end-to-end object detection with aligned loss

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.379980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:cc669714e47c03df23b7da8d76c163081bde8b88fa92f2f3b48ef206d745b781

Observation 93e86de1-a942-4a4d-a285-10faee621bd0 · outbound

This paper cites Observation-centric sort: Rethinking sort for robust multi-object tracking.

SAM 3: Segment Anything with Concepts Observation-centric sort: Rethinking sort for robust multi-object tracking

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.267527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:c88c7f8bdc177436ca5e851337fe7b9bc47367d384de0942cdd3a5b7c38b1300

Observation 1f7bd674-e994-43e3-afb6-59ed2f422ba4 · outbound

This paper cites End-to-end object detection with transformers.

SAM 3: Segment Anything with Concepts End-to-end object detection with transformers

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.368528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:304b391052be71a7fce9fd7a5347a519ee1e225c6e65eab065e58e0231a0d849

Observation 896aa9b6-6c1d-4d56-a6e2-844dd3e2287c · outbound

This paper cites LW-DETR: A Transformer Replacement to YOLO for Real-Time Detection.

SAM 3: Segment Anything with Concepts LW-DETR: A Transformer Replacement to YOLO for Real-Time Detection

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.523127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:ed343820df059b368043444b2bfe92e7a5e774c112bf581e15bbcfd788246857

Observation 1ce86dc3-e963-4488-8dd7-817603627b44 · outbound

This paper cites Sam4mllm: Enhance multi-modal large language model for referring expression segmentation.

SAM 3: Segment Anything with Concepts Sam4mllm: Enhance multi-modal large language model for referring expression segmentation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.264670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:bb7db1308781627f9039f3bb04c6dae69c5b9110d35a988042fd72da17f9e460

Observation 5b6836b1-e40e-45b9-8df0-66b4806832a0 · outbound

This paper cites Re-aligning language to visual objects with an agentic workflow.

SAM 3: Segment Anything with Concepts Re-aligning language to visual objects with an agentic workflow

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.366172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:87d3222d79f333f0d792e809f70a492b9da3d242bd08b3866b66f8ddbfea1faf

Observation ee75b244-db1b-4a98-b3eb-9e737e84aac9 · outbound

This paper cites Schwing, and Alexander Kirillov.

SAM 3: Segment Anything with Concepts Schwing, and Alexander Kirillov

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.370781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:6911b00a55ef3dc14bf09ecc471d63e4b4634be109561a2559c7db0f49d2cd07

Observation 9b872af4-c4ba-4404-9fd9-9132e9fe6cae · outbound

This paper cites PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding.

SAM 3: Segment Anything with Concepts PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.567661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:9f9dc5d0d7a66268a8c0c0d994668042804b136b7a0fac00cedbf97ce157cda7

Observation 9f83b8a9-4cfb-4b17-8a0c-76d1acb456cc · outbound

This paper cites ELECTRA : Pre-training text encoders as discriminators rather than generators.

SAM 3: Segment Anything with Concepts ELECTRA : Pre-training text encoders as discriminators rather than generators

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.394649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:78a9932c68888e511a504cff2bfa3cacdcf6b0b082209527ca6d87e2ea90b84d

Observation af39eeb1-05bd-4021-a9ba-e91c9578a9d6 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

SAM 3: Segment Anything with Concepts Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:25:11.445429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:7686f00ba4509e49bcf036333cccf7e54990393232c714fe6e1755098abdc8a2

Observation 362b6e69-149a-4c61-a79e-b87270247b25 · outbound

This paper cites The cityscapes dataset for semantic urban scene understanding.

SAM 3: Segment Anything with Concepts The cityscapes dataset for semantic urban scene understanding

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.504006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:b24ea3fdc310a50daefc5808f42ddef7a0e200495ac800b2feef6e6ed5a4de6c

Observation 1c0563c3-0eef-463a-bbab-44b536af04da · outbound

This paper cites Evaluating Large-Vocabulary Object Detectors: The Devil is in the Details.

SAM 3: Segment Anything with Concepts Evaluating Large-Vocabulary Object Detectors: The Devil is in the Details

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.431192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:dcf05ca7d1097fe5c2b0d1ac95d2fc16a59ff1e26d36fe5089a358eddb646c2c

Observation 627f38f0-73bb-4bd9-aa7a-a851a3934ff2 · outbound

This paper cites Molmo and pixmo: Open weights and open data for state-of-the-art vision-language models.

SAM 3: Segment Anything with Concepts Molmo and pixmo: Open weights and open data for state-of-the-art vision-language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.353805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:c1e6afe061ef0cb38eab1c4eade7edc68cfdfdd9826e3074e5d39440c3853087

Observation f7264151-16c5-4eac-bce8-4a8c3b6eab47 · outbound

This paper cites MOSEv2: A more challenging dataset for video object segmentation in complex scenes.

SAM 3: Segment Anything with Concepts MOSEv2: A more challenging dataset for video object segmentation in complex scenes

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.492824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:a4a69a8127ea1480f4a8253a802407d87a55b8e9e96f9f1b4dea905ea07af0b1

Observation 583d7e1f-081e-4836-8b46-80cfef7a4eb4 · outbound

This paper cites A large-scale synthetic pathological dataset for deep learning-enabled segmentation of breast cancer.

SAM 3: Segment Anything with Concepts A large-scale synthetic pathological dataset for deep learning-enabled segmentation of breast cancer

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.336699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:c9810cde1c958af8394041a025e4e2620df806d46d5ed4a614b1c5f40a7d3f11

Observation 554a66a4-47b5-4ab3-972c-077b837a3b5e · outbound

This paper cites SAM2Long: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory Tree.

SAM 3: Segment Anything with Concepts SAM2Long: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory Tree

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.455972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:c3293ac7008118a33da65450de4c95d31de3ea3e7f3c0d5710491bd4b0fef4d0

Observation 35a756b1-9635-468d-9d7b-f68ac69fd194 · outbound

This paper cites Open-Vocabulary Universal Image Segmentation with MaskCLIP.

SAM 3: Segment Anything with Concepts Open-Vocabulary Universal Image Segmentation with MaskCLIP

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.372628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:a43889872a85abf0b874e72f18df7c64530fcde0e2276f11d31f9c90bf01a3ea

Observation 6b3b6b65-7403-46f3-bc16-68e45c254392 · outbound

This paper cites Coarse-to-Fine Vision-Language Pre-training with Fusion in the Backbone.

SAM 3: Segment Anything with Concepts Coarse-to-Fine Vision-Language Pre-training with Fusion in the Backbone

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.547619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:fdf1be671453414b894db8df89e99c6c3898c78788ea6e0a415d1e32b929377d

Observation 0520de45-545c-4b27-8c35-5c99a2ccbba0 · outbound

This paper cites The llama 3 herd of models.

SAM 3: Segment Anything with Concepts The llama 3 herd of models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.331965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:a4ef29069533d528f95f277fe21a66fa5ab3724f2c059cd3eda24d8fcbcee4ba

Observation ed25477c-7854-4e8c-9377-8f6fe279cdd2 · outbound

This paper cites Livecell—a large-scale dataset for label-free live cell segmentation.

SAM 3: Segment Anything with Concepts Livecell—a large-scale dataset for label-free live cell segmentation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.361465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:2e70165917990b889936217f55481eaec30aea563b4d245d627e0761b617cff7

Observation 5a7738df-9ac7-4812-91ad-c42bc6cc4caf · outbound

This paper cites Detect to track and track to detect.

SAM 3: Segment Anything with Concepts Detect to track and track to detect

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.327049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:8c654d276c7f5a6b9dad4aeda72f4b5bc76c15d41208a36fd2d2063b109f1dbf

Observation 587a2805-146d-42db-a560-871b140a842b · outbound

This paper cites an unresolved cited work.

SAM 3: Segment Anything with Concepts Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:25:12.334280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:c3fe2eecee59b644c8c8bfc5fa3c347f5d815521f70f69baa367a46b27929423

Observation f3ab3d75-285a-4ec9-806c-df92ec7ad3ec · outbound

This paper cites Llmdet: Learning strong open-vocabulary object detectors under the supervision of large language models.

SAM 3: Segment Anything with Concepts Llmdet: Learning strong open-vocabulary object detectors under the supervision of large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.341566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:72594e636eb1a0bde17adafc480255049136acf4fb3d884a86f4e4b9cef16da7

Observation 4688bd17-a57f-4e42-a9d6-d3342f4343cb · outbound

This paper cites Pannuke: an open pan-cancer histology dataset for nuclei instance segmentation and classification.

SAM 3: Segment Anything with Concepts Pannuke: an open pan-cancer histology dataset for nuclei instance segmentation and classification

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.359197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:772bdcf4e9d902ff8d7da9e1034d9a7da195433bc9a926dba77cad7275cc79da

Observation 7f2a7a4e-512f-4469-ba0d-a9bb15fb0cf4 · outbound

This paper cites PanNuke Dataset Extension, Insights and Baselines.

SAM 3: Segment Anything with Concepts PanNuke Dataset Extension, Insights and Baselines

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:25:11.411152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:8d236529f1ead1bc7fb7eecaecf8e8184ecd60ec05f202e4161020d17b5ef470

Observation e01f6412-018d-4dda-8069-0b69a99a2a77 · outbound

This paper cites an unresolved cited work.

SAM 3: Segment Anything with Concepts Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:25:12.317721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:618fdd271a9864530af492499ee5f2f3e58689383f4065bbba52abe6c2552391

Observation 48f4fb16-2244-44e7-9240-93aa5bf2ddb4 · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

SAM 3: Segment Anything with Concepts Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:25:11.414343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:279406021375e4b7a067da23888b734f240066894cbe011708a78f7d0427ad42

Observation 98d8d0b7-3353-4487-8414-b87483717b75 · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

SAM 3: Segment Anything with Concepts Lvis: A dataset for large vocabulary instance segmentation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.322438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:32b89b23e79ca94fb7363117bccf99337e876c0161d99d0d5481150f10028230

Observation b3086335-c28b-4a49-ab4c-806d3097b412 · outbound

This paper cites Masked autoencoders are scalable vision learners.

SAM 3: Segment Anything with Concepts Masked autoencoders are scalable vision learners

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.305113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:e643ba053f5cf2a245c05d5cda3cf88b554fb80cde749f0227683ae5288de605

Observation 7e5377c8-8e90-4d66-99c8-52d0b7d89111 · outbound

This paper cites Rotary Position Embedding for Vision Transformer.

SAM 3: Segment Anything with Concepts Rotary Position Embedding for Vision Transformer

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.403942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:50e47cffeb9a8e4857dc372d5e305d35584bb9825c710cd94668bf8c81890254

Observation d3d4a90f-82dc-4e05-b183-0fa5a395fbdc · outbound

This paper cites LVOS: A Benchmark for Large-scale Long-term Video Object Segmentation.

SAM 3: Segment Anything with Concepts LVOS: A Benchmark for Large-scale Long-term Video Object Segmentation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.554163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:a05892465ca4bde08c402de1044765133af8d6d3229f3d550872e2f65d058f4f

Observation 6eb8893b-5295-49d8-b31d-0131b2258318 · outbound

This paper cites author Montani, I.

SAM 3: Segment Anything with Concepts author Montani, I

Reference 44

Resolution
metadata mismatch
doi, observed 2026-05-17T20:25:11.211822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:c96fcfc9bb33df50a81729d728a44f601ec1ca29adf01dc15c78c7058a780bf5

Observation f8f33147-a929-4cd6-bba5-99402320d17f · outbound

This paper cites The iNaturalist Species Classification and Detection Dataset.

SAM 3: Segment Anything with Concepts The iNaturalist Species Classification and Detection Dataset

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:25:11.387280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:f8e9b03b068d848ef48fb8a2f26fdeb6c5d2cb77b9cd7a1994ef0487fa725554

Observation 873da0d1-28ce-4939-88b1-d602b1d22d33 · outbound

This paper cites DAC-DETR : Divide the attention layers and conquer.

SAM 3: Segment Anything with Concepts DAC-DETR : Divide the attention layers and conquer

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.275119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:81cfe4cec201b085d56145cb39c7c7ccdf1d642f36e8d98f9da44fcacdd21546

Observation 788ce252-4f03-460e-b466-5e2311b1a7d3 · outbound

This paper cites Densely connected parameter-efficient tuning for referring image segmentation.

SAM 3: Segment Anything with Concepts Densely connected parameter-efficient tuning for referring image segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.277772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:d740d9902949ac14ac09d0cbc111ed62358f753a82c5cc4f7b189d37941131a6

Observation 83f811d4-57be-402f-ae80-99729362d600 · outbound

This paper cites DETRs with Hybrid Matching.

SAM 3: Segment Anything with Concepts DETRs with Hybrid Matching

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.393789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:dcd5dd6875d68f78239f6359d44e131b9f2ecd70abd9056bd3a1dccaf913e771

Observation 9e67fa51-2b6c-4f95-8a2b-531112388380 · outbound

This paper cites Fashionpedia: Ontology, Segmentation, and an Attribute Localization Dataset.

SAM 3: Segment Anything with Concepts Fashionpedia: Ontology, Segmentation, and an Attribute Localization Dataset

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.564043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:d5aa97cfaac54818190dbd338742cb7e14a56e4d754a0460f0371061a5fab9d5

Observation e218e163-f96b-4d89-a6c9-78a09ef003bd · outbound

This paper cites Sam2mot: A novel paradigm of multi-object tracking by segmentation.

SAM 3: Segment Anything with Concepts Sam2mot: A novel paradigm of multi-object tracking by segmentation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.427667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:ae8be8beed14ec25ac8bf4de1c3a8399957b14984e747dfc18a2c2991af66a83

Observation 3152d228-2a7f-4164-8087-beb56eeb75bf · outbound

This paper cites T-rex2: Towards generic object detection via text-visual prompt synergy.

SAM 3: Segment Anything with Concepts T-rex2: Towards generic object detection via text-visual prompt synergy

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.282360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:5c5f927cf43585b2eb047b9fed0a84ead2a0f7176e444645789b64f580dc8468

Observation a8cac499-bd4c-413d-91a6-a35043d95ce1 · outbound

This paper cites Trackeval.

SAM 3: Segment Anything with Concepts Trackeval

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.272382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:81250153fd6c7ad1e54a6411a91029716a3a0b74056165d81f1a3f3b1d526852

Observation 09ea6adb-25e3-4a50-aad8-cd84ca1c8415 · outbound

This paper cites Mdetr-modulated detection for end-to-end multi-modal understanding.

SAM 3: Segment Anything with Concepts Mdetr-modulated detection for end-to-end multi-modal understanding

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.501610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:77358e6652a8f253c2fd8fa7c2e299b6a4345466b2bde5c056dc304c8988c01b

Observation 4de276d2-9682-45f1-9322-b2fd161654e0 · outbound

This paper cites Your large vision-language model only needs a few attention heads for visual grounding.

SAM 3: Segment Anything with Concepts Your large vision-language model only needs a few attention heads for visual grounding

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.509076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:844994412b43de37c110a4850f6a534d9e6f380a91833a8464a97deaee53af68

Observation ce1242b9-1664-4dd6-96f0-66606f860f1e · outbound

This paper cites FathomNet: A global image database for enabling artificial intelligence in the ocean.

SAM 3: Segment Anything with Concepts FathomNet: A global image database for enabling artificial intelligence in the ocean

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.452255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:3c8d53239487a08a43e8edbb3f34b8e9639faa8386d792d348f1bc6f56ebb1d6

Observation afe9653d-60bb-4f74-85e7-f6c8c3823c7c · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

SAM 3: Segment Anything with Concepts Referitgame: Referring to objects in photographs of natural scenes

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.522938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:606eb3a42e9ec04761705f12ae1bc914102696b1b5f317607624043e5b1ab296

Observation d2aced7b-0cc2-4614-a2da-b97f731320e4 · outbound

This paper cites Video mask transfiner for high-quality video instance segmentation.

SAM 3: Segment Anything with Concepts Video mask transfiner for high-quality video instance segmentation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.525138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:33598b1f8aae1c100deff04457852e656b5491f9efe7415bc705ae2f46996d07

Observation 25413d31-9899-4168-8023-3c5d5f5d607f · outbound

This paper cites an unresolved cited work.

SAM 3: Segment Anything with Concepts Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:25:12.284584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:b7c2b3e9b9c574c67296acaebdb1053b1b266784551587fceba4132206a1e72e

Observation 099940c9-e8a8-43ea-b9ca-db73581b73a9 · outbound

This paper cites Sapiens: Foundation for Human Vision Models.

SAM 3: Segment Anything with Concepts Sapiens: Foundation for Human Vision Models

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:25:11.397091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:bc1ac781e328953552e072b1004cecdb47388566ce03395bf73d0ff04be6feb0

Observation 7842e0dd-ce7f-42c8-b965-14f56feb827e · outbound

This paper cites Segment anything.

SAM 3: Segment Anything with Concepts Segment anything

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.520673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:5a1dbf479c1a98cf441104c98dacb6afa20939f287be029e09546b8966c60d93

Observation cd4630c6-c9e6-47db-95c6-8108338b864c · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

SAM 3: Segment Anything with Concepts Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.488089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:f2a0ea159d4dd51a4a1b5679b5bd87e0e72f2d27ac93f42b6327062b57442d72

Observation dd70c8d3-d410-47c9-9600-3f5d16fdfb0c · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

SAM 3: Segment Anything with Concepts The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.485512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:37985dab5f6b111c21e545c9ba49cf33981c9077c45d58f911ec5914b05f53d0

Observation bec7a059-59f9-46ec-9ffd-f790b2dad35d · outbound

This paper cites Quantifying the Carbon Emissions of Machine Learning.

SAM 3: Segment Anything with Concepts Quantifying the Carbon Emissions of Machine Learning

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:25:11.533141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:8b87278423d883f438733a30344125800620f08eb83910d30fe0b1f512f87f27

Observation fd775d3b-89cc-48b0-afa8-04208920e2cd · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

SAM 3: Segment Anything with Concepts Lisa: Reasoning segmentation via large language model

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.480111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:3f4e0f48c1afdee4add32409d2ce3ccd826847684cc6c5ba6dda08d6e9394a68

Observation a196dd1d-ca4a-410a-847e-bb692f7d71ce · outbound

This paper cites EDEN: Multimodal Synthetic Dataset of Enclosed garDEN Scenes.

SAM 3: Segment Anything with Concepts EDEN: Multimodal Synthetic Dataset of Enclosed garDEN Scenes

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.483043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:59a9b886e263d2f17f5e8b7043276adaecc6519423f7ad450b6adb2315938c64

Observation 88a82047-d6d1-4d2d-89d2-a54958474c81 · outbound

This paper cites Elevater: A benchmark and toolkit for evaluating language-augmented visual models.

SAM 3: Segment Anything with Concepts Elevater: A benchmark and toolkit for evaluating language-augmented visual models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.490448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:4a00540a2f91e0eb3b8518160e5dd6856fdb6b706a64f52aeb1fc32ba61ea47f

Observation 1b54f151-db12-45d8-8211-1d4fd855b7fa · outbound

This paper cites Visual in-context prompting.

SAM 3: Segment Anything with Concepts Visual in-context prompting

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.456912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:d539ee0c3a6e169a965cbd764bc1bf172ad89f983aa81246b720e13feedfb131

Observation ad5b1c4f-3d5a-458f-9e1b-771cd5d2e131 · outbound

This paper cites LGD: Leveraging Generative Descriptions for Zero-Shot Referring Image Segmentation.

SAM 3: Segment Anything with Concepts LGD: Leveraging Generative Descriptions for Zero-Shot Referring Image Segmentation

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.512700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:23417321f555e39f0047b5958b732b36be599f49a005e3893dea873aa17153fc

Observation 43f652ba-fd9c-42eb-b0b1-6c1f6b2836fb · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

SAM 3: Segment Anything with Concepts BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:25:11.526629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:9a3fbc71a49c340377014f9c4f8d303d8d69bd4ee04d6363912a26cf96f4031b

Observation be2e9b4c-5ab2-4862-83c4-5982e0401363 · outbound

This paper cites Desco: Learning object recognition with rich language descriptions.

SAM 3: Segment Anything with Concepts Desco: Learning object recognition with rich language descriptions

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.444256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:9541587e1c7e0cf9e1c6af875433bf9dd2fdf47b7e7088bcd38139c84ecd96a4

Observation c4aa71e5-931f-4dce-8cc4-07aa9d1e97c4 · outbound

This paper cites Grounded language-image pre-training.

SAM 3: Segment Anything with Concepts Grounded language-image pre-training

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.448988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:d2482e4f711ac20dd34723f0fafad25bbaaf6b7f8a0d252b954645d0f80d591d

Observation fb425089-a84b-43ad-93e7-26770477cf52 · outbound

This paper cites Tracking every thing in the wild.

SAM 3: Segment Anything with Concepts Tracking every thing in the wild

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.451329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:7a122e6baa3abe5cc3e53e128e00393e3cc743cdd6714de7dedd91dd14e81c81

Observation 0145bc80-2102-4375-b791-279d6e4253de · outbound

This paper cites Exploring plain vision transformer backbones for object detection.

SAM 3: Segment Anything with Concepts Exploring plain vision transformer backbones for object detection

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.462655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:558d1da109a49214015afe181ed3f69e55dd98d3569818624427e4d396d13e45

Observation 54b4156a-7465-4598-bb8c-91ea648ed839 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

SAM 3: Segment Anything with Concepts Open-vocabulary semantic segmentation with mask-adapted clip

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.498861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:6b882c8647bebc8d0d68eb6558d01dc12981714dfe912980b0b74130e35e1dd3

Observation fd2cbe59-9dbc-4e9b-aa1b-ad523c3a66b4 · outbound

This paper cites WCS camera traps.

SAM 3: Segment Anything with Concepts WCS camera traps

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.515829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:297794dfa094610d3d6e877719ab8afe0b5e02c6b15019c0e7e94d96f0943143

Observation 45f56976-187f-4dcf-8e61-18c0d1d2bafe · outbound

This paper cites Microsoft coco: Common objects in context.

SAM 3: Segment Anything with Concepts Microsoft coco: Common objects in context

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.434784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:3ca29f134145332134636c1da2d009c21ce0597b04b67f317e95c90ce3dab54c

Observation 106a1d16-3f00-4810-a4a6-a2d6ff0c8cc9 · outbound

This paper cites Detr doesn't need multi-scale or locality design.

SAM 3: Segment Anything with Concepts Detr doesn't need multi-scale or locality design

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.425347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:f1d7c29fae1822594563464bbb8073ff71d7e65e92d242454e141b97f1c6eb41

Observation 2c219341-cf14-4b04-8196-3c456d3d4c62 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

SAM 3: Segment Anything with Concepts Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.427753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:ca382795503cdbff4a123c9b9de9692503ac360d6ef862cda39854b3e4873ab2

Observation 0ff77ce1-269d-44a0-8e0a-388ed9268477 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

SAM 3: Segment Anything with Concepts Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.430218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:c84ae6dfc9736370afc96879e234a5a26c235db6f1d61d6f1d27d6b3f4470a93

Observation bed37865-5700-40ed-b380-841760ae1b3a · outbound

This paper cites Hybrid global-local representation with augmented spatial guidance for zero-shot referring image segmentation.

SAM 3: Segment Anything with Concepts Hybrid global-local representation with augmented spatial guidance for zero-shot referring image segmentation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.418584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:1a7abfe0be338803e4f885c8bbbdaf7d400268ab28f502ea17559081877d1777

Observation 53dd0bad-7b08-4b4a-8131-9458ff9d9389 · outbound

This paper cites Universal segmentation at arbitrary granularity with language instruction.

SAM 3: Segment Anything with Concepts Universal segmentation at arbitrary granularity with language instruction

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.420702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:0272f40339b1a2bd82aaadace049515bf73aa587ad9ea0f7658cd4e60f010ee9

Observation 59c3125b-5bdd-4e2f-8258-eadd370c475c · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

SAM 3: Segment Anything with Concepts Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:25:11.470281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:4ed7a761487ed36e15bfcf68f20d46fb84778001a4119bbba5f522e48610270d

Observation 8b3f0a00-242b-4a55-a321-d35808c6a216 · outbound

This paper cites Towards end-to-end unified scene text detection and layout analysis.

SAM 3: Segment Anything with Concepts Towards end-to-end unified scene text detection and layout analysis

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.423179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:19ec07579e8e6566cf9bd2bc7df4030959c36f779929fc23e84bf7fe83f04d8f

Observation 2ac156fb-0aa7-456a-af71-47e6cbe6edaf · outbound

This paper cites ICDAR 2023 Competition on Hierarchical Text Detection and Recognition.

SAM 3: Segment Anything with Concepts ICDAR 2023 Competition on Hierarchical Text Detection and Recognition

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.489087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:79946c7ae4f37bd7e5616190d739448bdbe87fd712f3d8a78e91da7d381e2d87

Observation 6a6cd9ba-f971-4c22-8148-dac8d729c35f · outbound

This paper cites Decoupled weight decay regularization.

SAM 3: Segment Anything with Concepts Decoupled weight decay regularization

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.439532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:2842476269c4fd5ee7ffdb270824c791c28771ba34434c0809bd0ce290e6ef01

Observation 0ffd6ff2-f54f-4d8a-9d37-5055ffd9f570 · outbound

This paper cites RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought.

SAM 3: Segment Anything with Concepts RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.519580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:5e815ff9312b09770eacbc94ddec475d77cfa283d195ec0eb5da9aa9df1d9227

Observation fa41eb51-36bd-4fc8-b3c8-35912ed9ffb5 · outbound

This paper cites Hota: A higher order metric for evaluating multi-object tracking.

SAM 3: Segment Anything with Concepts Hota: A higher order metric for evaluating multi-object tracking

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.409327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:ab7cfa2ecb4afcf84b4991597fb9189416917f0b4ab36dfa4488843852cf9138

Observation aee0cbd5-fbcb-41ad-8e80-5b04981a4de0 · outbound

This paper cites MedSAM2: Segment Anything in 3D Medical Images and Videos.

SAM 3: Segment Anything with Concepts MedSAM2: Segment Anything in 3D Medical Images and Videos

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.441824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:abe99439d1f93e945b542ba842fe8a076d45d2723422856fb962cdaf7dfba7fe

Observation 500faeb0-3c71-46c2-ba6b-7102584321af · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

SAM 3: Segment Anything with Concepts Generation and comprehension of unambiguous object descriptions

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.411682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:18c052aa39a7b4aec44cf0c1c20d743f2b21a7ecf1eba100a6d8dfafb81ce48e

Observation 0a8a2c3c-2ab9-4f66-b15d-08f1e1a34b5e · outbound

This paper cites COCO-O: A Benchmark for Object Detectors under Natural Distribution Shifts.

SAM 3: Segment Anything with Concepts COCO-O: A Benchmark for Object Detectors under Natural Distribution Shifts

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.515971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:d36bb14a0d011bee8630e57ef35531078bf1df43e60b53f837b698061e3d8d4a

Observation 6fb45219-baf8-4d9a-96d0-acf12d4efc07 · outbound

This paper cites Trackformer: Multi-object tracking with transformers.

SAM 3: Segment Anything with Concepts Trackformer: Multi-object tracking with transformers

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.406571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:09f4906c637d903959f9d2650b508d01c93039833d57d4923129b5fbde64be08

Observation d29dc5ac-1351-4538-bded-c3a70212d3d0 · outbound

This paper cites Simple open-vocabulary object detection.

SAM 3: Segment Anything with Concepts Simple open-vocabulary object detection

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.413872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:dfeffaaf49b90df37f6a3363ef51316b7c0ddfa6dacbe463234ffae23e866d9f

Observation 75978a72-45bf-4ecc-bbee-a5079baaf29a · outbound

This paper cites Scaling Open-Vocabulary Object Detection.

SAM 3: Segment Anything with Concepts Scaling Open-Vocabulary Object Detection

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.477247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:125e4e3a245c17830a5fb8e92cf314aea65d6fd2e692e5aa1449c8fecf99f869

Observation 7ce729b6-c3a1-4872-a66e-9dbb725684ed · outbound

This paper cites ARMBench: An Object-centric Benchmark Dataset for Robotic Manipulation.

SAM 3: Segment Anything with Concepts ARMBench: An Object-centric Benchmark Dataset for Robotic Manipulation

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.508856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:cb7e24845a094afe1f61ddaea0d5fb4324bf0020cac80f58c0e7d0dc01b5a27f

Observation b09c948e-7c6a-4697-b790-99918bc65632 · outbound

This paper cites Model cards for model reporting.

SAM 3: Segment Anything with Concepts Model cards for model reporting

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.518406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:e913e61f694f5e59e1511ba206ed7ef6a730e9c83f2610c2c9a85e16569e9731

Observation d4417618-3315-4d4f-8d99-70b93171f163 · outbound

This paper cites The Food Recognition Benchmark: Using DeepLearning to Recognize Food on Images.

SAM 3: Segment Anything with Concepts The Food Recognition Benchmark: Using DeepLearning to Recognize Food on Images

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.544092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:62a1befc1ca0f24e9edacb31c88c9ed51a42af46196271f9d4357d91c43d32a4

Observation c0817322-7fdd-458c-b6ed-2bf12fabc242 · outbound

This paper cites The role of context for object detection and semantic segmentation in the wild.

SAM 3: Segment Anything with Concepts The role of context for object detection and semantic segmentation in the wild

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.534751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:88cc95c682ebab7bb5194eef52369122336f2efb5330a9f4710718ea1b467c0f

Observation 889241bd-e00f-4ac2-8199-9202dba4a043 · outbound

This paper cites Public domain collection dataset.

SAM 3: Segment Anything with Concepts Public domain collection dataset

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.432593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:11cbeee5b8feba2748afc23b09d6075f7759dbd6165a3cafe1b1eeffc9fde524

Observation 17ab4453-da85-4878-8c7a-1465cec38302 · outbound

This paper cites Ref-Diff: Zero-shot Referring Image Segmentation with Generative Models.

SAM 3: Segment Anything with Concepts Ref-Diff: Zero-shot Referring Image Segmentation with Generative Models

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.540542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:e4c68cfa9271a6a5d52c2f9fcfaeaa7d52eb5c19d8cab8d5672d2f8384212d80

Observation ee35e022-09f1-4f64-9675-c87e707ec7ff · outbound

This paper cites Dinov2: Learning robust visual features without supervision.

SAM 3: Segment Anything with Concepts Dinov2: Learning robust visual features without supervision

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:25:12.495940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:1126ebb51278783de84009d68f76de7210bf902f7da66095996adebe37110237

Pith citing papers

Observation c7d6612a-9866-4d67-aad0-d2e308e7b740 · inbound

Glass Surface Detection: Leveraging Reflection Dynamics in Flash/No-flash Imagery cites this paper.

Glass Surface Detection: Leveraging Reflection Dynamics in Flash/No-flash Imagery SAM 3: Segment Anything with Concepts

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:03.793460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:03.793460Z digest=sha256:01e4c9488851b168418970a632a16198dec19e2c26af78b8671c1a0415b0897f

Observation a2a1316d-48e4-4cb7-a584-3c9f1cef3bb1 · inbound

Comparing SAM 2 and SAM 3 for Zero-Shot Segmentation of 3D Medical Data cites this paper.

Comparing SAM 2 and SAM 3 for Zero-Shot Segmentation of 3D Medical Data SAM 3: Segment Anything with Concepts

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T04:39:02.640472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T04:38:15.388826Z digest=sha256:085e820e265c12e85faf9cbf6a3b75f027b45b8fd07da6a9b1d60329eb119067

Observation 78513a55-c14a-4302-95ec-dbfebe3777ca · inbound

SAM3-I: Segment Anything with Instructions cites this paper.

SAM3-I: Segment Anything with Instructions SAM 3: Segment Anything with Concepts

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:08:51.402283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T02:06:35.864921Z digest=sha256:02c7434269fce19fe2ddabdc3085f242a79b0118070e7250e0965d2ac509491b

Observation 8915b35e-d61f-45d5-8a7d-37a833b1c4ee · inbound

Automated Annotation of Shearographic Measurements Enabling Weakly Supervised Defect Detection cites this paper.

Automated Annotation of Shearographic Measurements Enabling Weakly Supervised Defect Detection SAM 3: Segment Anything with Concepts

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:21:23.687542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:19:31.834944Z digest=sha256:aad2adf39dfbf8139b3177012d465b9f1d00c756ec897c409402428064b27fb4

Observation a0adf699-8f55-4589-85d0-623d4cbbecbf · inbound

SegEarth-OV3: Exploring SAM 3 for Open-Vocabulary Semantic Segmentation in Remote Sensing Images cites this paper.

SegEarth-OV3: Exploring SAM 3 for Open-Vocabulary Semantic Segmentation in Remote Sensing Images SAM 3: Segment Anything with Concepts

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:53:42.289506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T23:53:27.350335Z digest=sha256:1fdad10eed174b1d4934912a282d3968016ab722bee57f4172d388df05ef2281

Observation d77353ef-1ffc-4652-80d9-0ec8b0696e77 · inbound

EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal cites this paper.

EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal SAM 3: Segment Anything with Concepts

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T14:08:20.175019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:08:20.175019Z digest=sha256:684b002030ab67bc60467675d6ee24150940b6f494a7f05f2289676506290a26

Observation 72b453bb-ccd7-4704-a95c-eee46a382203 · inbound

PanoSAMic: Panoramic Image Segmentation from SAM Feature Encoding and Dual View Fusion cites this paper.

PanoSAMic: Panoramic Image Segmentation from SAM Feature Encoding and Dual View Fusion SAM 3: Segment Anything with Concepts

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T14:48:00.841351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T14:43:18.682997Z digest=sha256:302c4a3f4fdb941921462e3ae08fbc447354fbf0c856bdd98b2c059a53616fe9

Observation 4965ad96-e67f-4501-bbf1-aceb171e428a · inbound

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding cites this paper.

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding SAM 3: Segment Anything with Concepts

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-16T04:21:29.693226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T04:21:29.526008Z digest=sha256:46b163824a0ce4d0059a67c109b8be21215c4534eeaf246a1058cb403c5945f3

Observation 5eb98aaf-4d8f-456b-aa38-e587bcff05f0 · inbound

OmniOVCD: Streamlining Open-Vocabulary Change Detection with SAM 3 cites this paper.

OmniOVCD: Streamlining Open-Vocabulary Change Detection with SAM 3 SAM 3: Segment Anything with Concepts

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:52:53.290843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T12:51:19.006683Z digest=sha256:c72b6c614d00b8783b2b125d7baddafaf99a02fe91ee85a2235c20033f2df839

Observation ae6b334e-9778-49bb-9120-fdc5985a915f · inbound

Going with the Flow: Koopman Behavioral Models as Pseudo Planners for Visuo-Motor Dexterity cites this paper.

Going with the Flow: Koopman Behavioral Models as Pseudo Planners for Visuo-Motor Dexterity SAM 3: Segment Anything with Concepts

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T03:41:53.322999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:41:53.322999Z digest=sha256:c366ecfe65c9e7e49a9b0e16011de5123d6442838d57be133465b01d173af464

Observation 081f0df1-d124-4cc9-abe0-b3032db15574 · inbound

DenseMLLM: Standard Multimodal LLMs for Dense Prediction cites this paper.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction SAM 3: Segment Anything with Concepts

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:33.793070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:33.793070Z digest=sha256:65cc916f2848d1c22daee1a634259925f145c4e24847772157a42779b7db60a2

Observation 09ff3779-20df-43e0-b28e-45d8bef4a2f0 · inbound

CAD-Prompted SAM3: Geometry-Conditioned Instance Segmentation for Industrial Objects cites this paper.

CAD-Prompted SAM3: Geometry-Conditioned Instance Segmentation for Industrial Objects SAM 3: Segment Anything with Concepts

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T21:20:16.412077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:20:16.412077Z digest=sha256:16b689e7dae6f5dca0335525306e982bdc3636fe8d82ae026f5319deed49d2d7

Observation 9feaafd4-f012-4749-8bf5-90300f24e668 · inbound

From Local Matches to Global Masks: Template-Guided Instance Detection and Segmentation in Open-World Scenes cites this paper.

From Local Matches to Global Masks: Template-Guided Instance Detection and Segmentation in Open-World Scenes SAM 3: Segment Anything with Concepts

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-15T16:20:09.762625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T16:19:15.019645Z digest=sha256:11c7f9a5a6c5b3e1281decc12c6f751d32cfa8fae240a6126fb7057eabf09864

Observation db76e5a5-b39d-41d9-a1d3-6d8f21d6cbf8 · inbound

OpenFrontier: General Navigation with Visual-Language Grounded Frontiers cites this paper.

OpenFrontier: General Navigation with Visual-Language Grounded Frontiers SAM 3: Segment Anything with Concepts

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-21T11:40:03.479992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T11:36:24.334860Z digest=sha256:8321d3bfcf2211e26200422fdcda1d66a43fc149ded69ee3871f6e630b1b1a66

Observation 5c6c7158-acad-4860-b5b0-45c668acc6f1 · inbound

JOPP-3D: Joint Open Vocabulary Semantic Segmentation on Point Clouds and Panoramas cites this paper.

JOPP-3D: Joint Open Vocabulary Semantic Segmentation on Point Clouds and Panoramas SAM 3: Segment Anything with Concepts

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-15T13:59:22.441714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:59:22.441714Z digest=sha256:20e3e4b686e0ed8319d3b5289f1384dd7d4fa78c7a667684afcfe9cb9edfbea4

Observation 126b5957-dd78-473d-b437-8dcdd122f9b2 · inbound

OPTED: Open Preprocessed Trachoma Eye Dataset Using Zero-Shot SAM 3 Segmentation cites this paper.

OPTED: Open Preprocessed Trachoma Eye Dataset Using Zero-Shot SAM 3 Segmentation SAM 3: Segment Anything with Concepts

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-15T14:35:55.821273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T14:33:41.517193Z digest=sha256:c09d25a509aee164a6e767881890593698327d933c0f39059dccf7c46362f50a

Observation 94da9b96-4aa6-4a33-8b4f-6a230249df7d · inbound

Margin in Abstract Spaces cites this paper.

Margin in Abstract Spaces SAM 3: Segment Anything with Concepts

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-15T13:24:52.801459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:24:52.801459Z digest=sha256:fed772ea9a56615a7d418bcccce7051123181d88865d667a423bb45b92d31845

Observation 0b8d655f-7c74-4a8f-81f2-ffe1340474fc · inbound

TrianguLang: Geometry-Aware Semantic Consensus for Pose-Free 3D Localization cites this paper.

TrianguLang: Geometry-Aware Semantic Consensus for Pose-Free 3D Localization SAM 3: Segment Anything with Concepts

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T15:16:09.686640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T15:12:23.459579Z digest=sha256:4620e9f6d3c5597f5f95d550127ae7ac36fe04417b03406b562a415a8f97f2a1

Observation e49c95a7-8b16-4f41-bb02-ecc23d38fa61 · inbound

From Ideal to Real: Stable Video Object Removal under Imperfect Conditions cites this paper.

From Ideal to Real: Stable Video Object Removal under Imperfect Conditions SAM 3: Segment Anything with Concepts

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T13:35:51.535952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T13:34:45.782304Z digest=sha256:10579589a50fd1a7ec9ade32b0201234b0ac63378cb2b1fd3e4edce7e8ba1018

Observation 8eaae0ae-bbad-4dd4-b1fa-0a2fb56006ca · inbound

PicoSAM3: Real-Time In-Sensor Region-of-Interest Segmentation cites this paper.

PicoSAM3: Real-Time In-Sensor Region-of-Interest Segmentation SAM 3: Segment Anything with Concepts

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T22:32:36.129552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:32:36.129552Z digest=sha256:3d94e2ccb781d040a36ea85fc9c70ed202e3ceb211d231b7fe112a7699475afb

Observation 9bcbec49-edef-4cfb-b0a0-b68d235da30d · inbound

RoboStereo: Dual-Tower 4D Embodied World Models for Unified Policy Optimization cites this paper.

RoboStereo: Dual-Tower 4D Embodied World Models for Unified Policy Optimization SAM 3: Segment Anything with Concepts

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T12:15:34.455286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T12:13:50.985113Z digest=sha256:f8edbd286e8b9abc2383717e32fed884b8c2f73fdfccd7aa91d1f3d26c3ff76a

Observation 092e52dc-08d3-44e7-9df4-d23ba99ce77e · inbound

RoboStream: Weaving Spatio-Temporal Reasoning with Memory in Vision-Language Models for Robotics cites this paper.

RoboStream: Weaving Spatio-Temporal Reasoning with Memory in Vision-Language Models for Robotics SAM 3: Segment Anything with Concepts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T21:59:30.885343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T21:59:30.885343Z digest=sha256:f868d9d1ec0f3bdfd13f836c3c830c106a2805aa5e71701ce8d184371798b837

Observation c32198cc-c6e5-4956-b528-4d10afc70f62 · inbound

Power Term Polynomial Algebra for Boolean Logic cites this paper.

Power Term Polynomial Algebra for Boolean Logic SAM 3: Segment Anything with Concepts

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T21:36:54.932072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T21:36:54.932072Z digest=sha256:222a3c2f55c6a0b16551121dd865ac464d3b403fcd6194488ffcf7f22d87229f

Observation 51eaa2b5-760c-45ac-bd63-64c39a805556 · inbound

SegviGen: Repurposing 3D Generative Model for Part Segmentation cites this paper.

SegviGen: Repurposing 3D Generative Model for Part Segmentation SAM 3: Segment Anything with Concepts

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:29:53.456812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T09:29:08.513396Z digest=sha256:eb262c08a38ce572217a0fb2b54a8ff9e08cd96494220b5658c6d9a5c4b02c28

Observation 42ab5947-cc6c-4164-90b3-5b8acf0b23a1 · inbound

AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors cites this paper.

AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors SAM 3: Segment Anything with Concepts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T22:49:03.259461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:49:03.259461Z digest=sha256:38a9eb9ca2747ef612c6b1220e1dfb64cef8dbfdeaf5fdfc2bd70af95af4766f

Observation e5f99ae3-0b2f-492f-9fe1-4a862a441a80 · inbound

Quantum orientation entanglement analysis of the interpolating helicity states between the instant form dynamics and the light-front dynamics cites this paper.

Quantum orientation entanglement analysis of the interpolating helicity states between the instant form dynamics and the light-front dynamics SAM 3: Segment Anything with Concepts

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T22:46:37.452699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:46:37.452699Z digest=sha256:4095a62cb21ea8b3519ce745069fcb07c236abcbaacedc3dbc1cfb285de01336

Observation 7752e164-6d7c-4fe9-8b44-10694d87b156 · inbound

TSegAgent: Zero-Shot Tooth Segmentation via Geometry-Aware Vision-Language Agents cites this paper.

TSegAgent: Zero-Shot Tooth Segmentation via Geometry-Aware Vision-Language Agents SAM 3: Segment Anything with Concepts

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:05:20.301673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T09:02:41.561003Z digest=sha256:4c9b2ede3968b738cd2f55d15877ca344c4a6b1fd5ba6a914628cd9a2395965b

Observation adf46d93-1c07-46fe-b065-32748c062e05 · inbound

Evaluating Vision Foundation Models for Pixel and Object Classification in Microscopy cites this paper.

Evaluating Vision Foundation Models for Pixel and Object Classification in Microscopy SAM 3: Segment Anything with Concepts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T17:52:42.442465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:52:42.442465Z digest=sha256:967ee9cd2ba762179788ab5eeb8d4832bab6136a157933ed382b68de9ff29c24

Observation 2aebbc26-3777-408f-8d1e-4c709306d49f · inbound

Memory Over Maps: 3D Object Localization Without Reconstruction cites this paper.

Memory Over Maps: 3D Object Localization Without Reconstruction SAM 3: Segment Anything with Concepts

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:49:50.970769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T07:46:50.155348Z digest=sha256:109d91a0f40769e84ba305e2c06a2f10c090ce0020a4b4572d718fa29faf80d5

Observation ba339898-8eff-4884-abb5-bddf055d02c1 · inbound

VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models cites this paper.

VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models SAM 3: Segment Anything with Concepts

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:59:36.760683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T00:54:36.125258Z digest=sha256:034e7ae6a7ad46341501e1f2a9d2b1704f006c33f941f3192331cccca41b47b2

Observation c0d0156a-a90a-48c1-ace4-14a837ab822c · inbound

DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection cites this paper.

DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection SAM 3: Segment Anything with Concepts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T19:38:38.452093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T19:38:38.452093Z digest=sha256:4ce54356e41a67684f1f5237514cccb0adbbf3998113fc749864cdc2098a543c

Observation 5472be30-f914-4499-a6cf-ec200f207c56 · inbound

ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis cites this paper.

ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis SAM 3: Segment Anything with Concepts

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T18:24:58.428701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:24:58.428701Z digest=sha256:fcfb1d06190c58d94f96e92ccc8148f1ee59ef35ac65028c1b67137acb19d1be

Observation 9638240a-360f-4a82-b14b-0369bfdeb041 · inbound

Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy Videos cites this paper.

Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy Videos SAM 3: Segment Anything with Concepts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T18:06:18.381075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:06:18.381075Z digest=sha256:dc1004930b99b4c6326b0bbbce42d59269ddd816c926cba84b84e611db2690da

Observation b2a0b6d2-0404-4899-820f-7c5e40e342b0 · inbound

MAR3: Multi-Agent Recognition, Reasoning, and Reflection for Reference Audio-Visual Segmentation cites this paper.

MAR3: Multi-Agent Recognition, Reasoning, and Reflection for Reference Audio-Visual Segmentation SAM 3: Segment Anything with Concepts

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-15T11:46:27.666803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T11:46:27.666803Z digest=sha256:c3dc13e7c68c540b4965727880146649710dc3e9053dc14167fb8a6b657a9ab7

Observation 4c51da11-f076-4283-b950-bd59cd1cfa9a · inbound

DRIVE-Nav: Directional Reasoning, Inspection, and Verification for Efficient Open-Vocabulary Navigation cites this paper.

DRIVE-Nav: Directional Reasoning, Inspection, and Verification for Efficient Open-Vocabulary Navigation SAM 3: Segment Anything with Concepts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T16:13:29.986339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:13:29.986339Z digest=sha256:bba59c3b54007e03c3be63850c8d0f35d48d929a2f44f7434d5aaaa43bf4f30c

Observation 9db1ee9b-c2d5-465d-b2ec-d45e6f9eb50c · inbound

EgoSim: Egocentric World Simulator for Embodied Interaction Generation cites this paper.

EgoSim: Egocentric World Simulator for Embodied Interaction Generation SAM 3: Segment Anything with Concepts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T14:40:24.818842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:40:24.818842Z digest=sha256:c339aeaa621a8491bf4496e695d570c6d18d5a871355ffa854d36453a5d219e6

Observation d54443b7-90a1-4b65-8230-c44d5aed02b5 · inbound

Rapidly deploying on-device eye tracking by distilling visual foundation models cites this paper.

Rapidly deploying on-device eye tracking by distilling visual foundation models SAM 3: Segment Anything with Concepts

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:43:19.083820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T21:38:55.454229Z digest=sha256:d53a6fda23f000af8d53279661c3f98008c412fe4fd653ddb328731c73dbb843

Observation ebfc333b-c4b7-4910-b9fb-54e1ec95ff0b · inbound

Moondream Segmentation: From Words to Masks cites this paper.

Moondream Segmentation: From Words to Masks SAM 3: Segment Anything with Concepts

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T20:18:13.720508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T20:14:45.629804Z digest=sha256:63b10aa30bec42ac0260475e8e9493cf990f59548730a238f02d5b5a45fc7ac5

Observation 7a048f2d-b3bd-4b28-b6c7-4248c5a52b06 · inbound

Generalized Small Object Detection:A Point-Prompted Paradigm and Benchmark cites this paper.

Generalized Small Object Detection:A Point-Prompted Paradigm and Benchmark SAM 3: Segment Anything with Concepts

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:13:13.562022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T20:11:01.013199Z digest=sha256:fc0eb977e8fed118e11b7432791aa0ea2c227cdcc77943e527d4f5b3ac91af7f

Observation da13772d-4786-412b-b9df-2e7d51ff0359 · inbound

KappaFormer: Physics-aware Transformer for lattice thermal conductivity via cross-domain transfer learning cites this paper.

KappaFormer: Physics-aware Transformer for lattice thermal conductivity via cross-domain transfer learning SAM 3: Segment Anything with Concepts

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-13T13:13:22.720572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T13:13:22.720572Z digest=sha256:486f02e8c483f8bc7219ae7495142de532c3e22c417baff7163cf95f0aa4c171

Observation e914de77-f439-4790-90f7-f5d68b1d84f0 · inbound

Beyond Standard Benchmarks: A Systematic Audit of Vision-Language Model's Robustness to Natural Semantic Variation Across Diverse Tasks cites this paper.

Beyond Standard Benchmarks: A Systematic Audit of Vision-Language Model's Robustness to Natural Semantic Variation Across Diverse Tasks SAM 3: Segment Anything with Concepts

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-10T21:55:52.598548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T20:26:42.128350Z digest=sha256:2857cada8db7d30a016662b4d85ade6e3c85e516c97cc288d543d88d24500b22

Observation 8ca6106a-f7f0-49e3-b8c6-73c28b58fa9a · inbound

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward cites this paper.

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward SAM 3: Segment Anything with Concepts

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:20:47.635927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:59:19.379119Z digest=sha256:810717b7d970f4b557b54de038a1529f4745e92b716d34fbac9312ed9df30f41

Observation f447b548-1169-41c7-ba3d-97cd4f91449e · inbound

SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing cites this paper.

SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing SAM 3: Segment Anything with Concepts

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T23:00:50.445165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:23:50.614589Z digest=sha256:6e2ae0f1f273748aef45c10b4f9084f372460940f686cd9356e0e292d0608f46

Observation 4a4c7d2c-9aad-415f-af00-b2ce866881a4 · inbound

MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation cites this paper.

MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation SAM 3: Segment Anything with Concepts

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-13T17:13:00.935260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T17:12:38.618233Z digest=sha256:fd5ae39314c46036e578ff9bf8c140d6d9fee2ead410d8f98b082c2059350836

Observation 42201890-872f-42b3-9d46-69034011ef9a · inbound

MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation cites this paper.

MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation SAM 3: Segment Anything with Concepts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T19:54:25.294067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:54:25.294067Z digest=sha256:b448853b0438cd6e15024a044381ff2aa98fee068e0c7c9ee8721977b7d978f9

Observation baa270bb-7c9a-46c3-bc6a-45b790afcd3f · inbound

Telescope: Learnable Hyperbolic Foveation for Ultra-Long-Range Object Detection cites this paper.

Telescope: Learnable Hyperbolic Foveation for Ultra-Long-Range Object Detection SAM 3: Segment Anything with Concepts

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T23:55:52.689492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:46:48.550865Z digest=sha256:86450be980e19d0b6db949de1783be1ceb3e4a282e2b63fe3a1990c9cbe12d27

Observation 4ac903d8-f713-4e18-9c55-8725f3500c0e · inbound

4D Vessel Reconstruction for Benchtop Thrombectomy Analysis cites this paper.

4D Vessel Reconstruction for Benchtop Thrombectomy Analysis SAM 3: Segment Anything with Concepts

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T20:25:46.968273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:16:38.725224Z digest=sha256:d8f1eafc137d2bbc9e98700f74b6a3d7345efc54935efb342cb0c75adb7e5494

Observation d9efa4a3-eb88-4be4-b8d0-cf4b27715efc · inbound

Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning cites this paper.

Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning SAM 3: Segment Anything with Concepts

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T23:15:47.800798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:16:46.753641Z digest=sha256:000b1ca7b6102ec06f0d4b2c1695224d192193a6a5bcb64e0b3a5e0bdf23bcb6

Observation 03c6a5e6-ec34-460a-b925-1e56076ae1cb · inbound

RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details cites this paper.

RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details SAM 3: Segment Anything with Concepts

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T23:20:54.386356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:11:43.172296Z digest=sha256:b0c10c9b50e43ee67374ac2bb2408319da3a86b233e18006e781fbd34e57b9db

Observation 5d5b4d84-0efc-4650-a5bb-0b9d7b8c2737 · inbound

PhyEdit: Towards Real-World Object Manipulation via Physically-Grounded Image Editing cites this paper.

PhyEdit: Towards Real-World Object Manipulation via Physically-Grounded Image Editing SAM 3: Segment Anything with Concepts

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:45:50.485034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:54:47.451103Z digest=sha256:fc7065391d6a7dc469d9c0d122aac74718a49c3b35f36109acccabab5bd5b194

Observation 60c2e017-28a6-4ccb-8bea-055e2ab0cee2 · inbound

Open-Ended Video Game Glitch Detection with Agentic Reasoning and Temporal Grounding cites this paper.

Open-Ended Video Game Glitch Detection with Agentic Reasoning and Temporal Grounding SAM 3: Segment Anything with Concepts

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:41:03.640206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:24:23.950768Z digest=sha256:77c65dfb958a859220c51c24d81ffbb50fa7f2850757602834dddb81b4ed9ba3

Observation d044e828-4fe9-421c-a63c-eaace89e72c8 · inbound

Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation cites this paper.

Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation SAM 3: Segment Anything with Concepts

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:21:00.963573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:11:13.376684Z digest=sha256:5202b8cc70753a8ae2330877b8cfbfe62a9deeb1e53e1137995a0af198821f9c

Observation 6d65d5da-0eeb-4d8b-baf6-0fe445301d1e · inbound

WildDet3D: Scaling Promptable 3D Detection in the Wild cites this paper.

WildDet3D: Scaling Promptable 3D Detection in the Wild SAM 3: Segment Anything with Concepts

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:26:00.661055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:38:13.336003Z digest=sha256:d065b5a4bfa7f1d5c2d7c3f9d19817cfca082290c6b65dbe330d1b8e7593d102

Observation fa03c5b0-d5af-403f-8b03-6252d42c85d4 · inbound

Generative Simulation for Policy Learning in Physical Human-Robot Interaction cites this paper.

Generative Simulation for Policy Learning in Physical Human-Robot Interaction SAM 3: Segment Anything with Concepts

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:35:57.540789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:06:58.099881Z digest=sha256:a7b34fee90c5dec171b81e19299faddde17734c634c713428abc017ea2952a5e

Observation 111f5887-71b7-455a-a96f-a6bf345e301b · inbound

Scene-Agnostic Object-Centric Representation Learning for 3D Gaussian Splatting cites this paper.

Scene-Agnostic Object-Centric Representation Learning for 3D Gaussian Splatting SAM 3: Segment Anything with Concepts

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:41:32.285461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:29:54.777890Z digest=sha256:4f088ba753c410b3c22899a530f4ca0b854580ba89dd26cc2836ae8fa3c9781a

Observation 7ab1e9fa-41d6-4c0a-af89-462ecc305336 · inbound

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding cites this paper.

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding SAM 3: Segment Anything with Concepts

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T06:51:23.259350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:25:31.097385Z digest=sha256:7cf8d0bc38390d98bfa4dc7f2d97c981adc182830ea2f907ab856508c68f34f6

Observation a13b7f91-19a5-4b3a-ac7f-43c4c9217348 · inbound

Are We Recognizing the Jaguar or Its Background? A Diagnostic Framework for Jaguar Re-Identification cites this paper.

Are We Recognizing the Jaguar or Its Background? A Diagnostic Framework for Jaguar Re-Identification SAM 3: Segment Anything with Concepts

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:20:48.897656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:55:56.261335Z digest=sha256:a097dc884ce45af599111032a5d97d491b060ffa3e27fda0dd5d83696ee6a6eb

Observation 168d0bfb-403b-44a3-834c-c1486b418785 · inbound

Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents cites this paper.

Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents SAM 3: Segment Anything with Concepts

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T07:51:01.853072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:55:48.506049Z digest=sha256:84aadb2d3d5649b7d61d7a6cee29429a1faf4beda4ab771834f96f9643acbe9b

Observation 503401ba-2aca-47fe-8edc-bbc80f3a4436 · inbound

Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents cites this paper.

Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents SAM 3: Segment Anything with Concepts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T23:07:40.699400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T23:07:40.699400Z digest=sha256:cd60c7d6645a4a01605935fdbceba5a4d7073c8781a3f1bc9ecec7f2f41757ca

Observation 36de14df-a458-47f3-8c76-e5127d64d09c · inbound

Does Your VFM Speak Plant? The Botanical Grammar of Vision Foundation Models for Object Detection cites this paper.

Does Your VFM Speak Plant? The Botanical Grammar of Vision Foundation Models for Object Detection SAM 3: Segment Anything with Concepts

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:35:57.207091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:07:17.174815Z digest=sha256:66d95231e48306942800c0041d60f0a0ebcb3452aa1c2e33179af209ab69c1e9

Observation 155740ca-7606-4443-9319-dddf1e39e451 · inbound

JARVIS: A Just-in-Time Augmented Reality VLM-Powered Instruction System for Cross-Reality Task Guidance cites this paper.

JARVIS: A Just-in-Time Augmented Reality VLM-Powered Instruction System for Cross-Reality Task Guidance SAM 3: Segment Anything with Concepts

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:00:59.364126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:21:13.585442Z digest=sha256:0de0ca5e6da5fe55e9a15f543a25c1f5a65a76ac62d1a903bb7ae7a845b32cb6

Observation e6aafc74-8893-4eea-8ee0-df58b60350e3 · inbound

JARVIS: A Just-in-Time Augmented Reality VLM-Powered Instruction System for Cross-Reality Task Guidance cites this paper.

JARVIS: A Just-in-Time Augmented Reality VLM-Powered Instruction System for Cross-Reality Task Guidance SAM 3: Segment Anything with Concepts

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:29:22.420008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T01:28:04.022431Z digest=sha256:18a0199b76797e3fd0b5f184940e20ff0cb769b65e3d603968dd7cc53f1e82f0

Observation f24a5d44-da82-4fd1-9018-6315e28ade09 · inbound

Semantic Manipulation Localization cites this paper.

Semantic Manipulation Localization SAM 3: Segment Anything with Concepts

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:21:01.491684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:32:30.345594Z digest=sha256:8d382d62b384c39df1781e993fd6e2d4830937f13414b1bebdf69c0c8938437c

Observation 8e1bdb60-defa-46c5-beb0-53e74c76113a · inbound

TInR: Exploring Tool-Internalized Reasoning in Large Language Models cites this paper.

TInR: Exploring Tool-Internalized Reasoning in Large Language Models SAM 3: Segment Anything with Concepts

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T22:20:17.569353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:20:17.569353Z digest=sha256:fab6172138684a4cb67230c695ffb2c725048deaaaa851011ce17c280e883f7f

Observation 7114e3ae-75e2-4503-8731-ef478c689045 · inbound

ReplicateAnyScene: Zero-Shot Video-to-3D Composition via Textual-Visual-Spatial Alignment cites this paper.

ReplicateAnyScene: Zero-Shot Video-to-3D Composition via Textual-Visual-Spatial Alignment SAM 3: Segment Anything with Concepts

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T10:06:05.406339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:38:03.377613Z digest=sha256:7c3a01096c70998d58b4ee8b16a73bd6177cb40d03d480e7a97b469be061873d

Observation 0ae249b4-71a2-4afb-9457-fa6e5d8886b6 · inbound

Do Instance Priors Help Weakly Supervised Semantic Segmentation? cites this paper.

Do Instance Priors Help Weakly Supervised Semantic Segmentation? SAM 3: Segment Anything with Concepts

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T16:30:36.253577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:64a538efa2ceff1b540df8eace031cdbda7081e46caa7679a02ddfeecae5f0e1

Observation 466f2ded-49f6-4c17-af7c-d20ee55e9d6c · inbound

Seg2Change: Adapting Open-Vocabulary Semantic Segmentation Model for Remote Sensing Change Detection cites this paper.

Seg2Change: Adapting Open-Vocabulary Semantic Segmentation Model for Remote Sensing Change Detection SAM 3: Segment Anything with Concepts

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:36:02.861601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:25:38.084033Z digest=sha256:ff0d5dc1559710b1ff17856a7e25e4cbd9807dafacb90639dfe6f73356150bc5

Observation 3c196fa8-8911-4584-b03b-baf8eaa14d18 · inbound

Online Reasoning Video Object Segmentation cites this paper.

Online Reasoning Video Object Segmentation SAM 3: Segment Anything with Concepts

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T10:31:00.229266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:03.843441Z digest=sha256:fdcde78d5518a5f74aa0af5246ae6f224a68837bfd71e596f91445ce79ee4101

Observation a152bce4-0965-4234-8347-7b5e03f39abb · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation SAM 3: Segment Anything with Concepts

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:11:04.723038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:0ddf3f9a4d82275b21068a7b06c09d4c36756f2e3571deef2250ecdf8b2930bb

Observation 47ed6040-9e11-43fd-8c76-a79a40a03696 · inbound

The Second Challenge on Cross-Domain Few-Shot Object Detection at NTIRE 2026: Methods and Results cites this paper.

The Second Challenge on Cross-Domain Few-Shot Object Detection at NTIRE 2026: Methods and Results SAM 3: Segment Anything with Concepts

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:21:00.673605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:05:16.070144Z digest=sha256:a9ca989b24913fab017e58800aba5c428feb9ee780410790f572e867aca6ae3a

Observation 648187dc-b0eb-4d1e-8f4b-9360cb3fb12e · inbound

VERITAS: A Multi-Agent Co-Scientist for Verifiable Image-Derived Hypothesis Testing cites this paper.

VERITAS: A Multi-Agent Co-Scientist for Verifiable Image-Derived Hypothesis Testing SAM 3: Segment Anything with Concepts

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:26:02.154005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T14:56:06.372343Z digest=sha256:960d2025fdb8c57f9a5acacb064565faddde8fd6824d68420e75370f5f20c1a3

Observation cba463f3-cbad-47d6-ab9e-86ce111eed04 · inbound

Grasp in Gaussians: Fast Monocular Reconstruction of Dynamic Hand-Object Interactions cites this paper.

Grasp in Gaussians: Fast Monocular Reconstruction of Dynamic Hand-Object Interactions SAM 3: Segment Anything with Concepts

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T11:21:00.599214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:01:52.354836Z digest=sha256:89b4967ba5d6cc730a0d8cf06d781e12505ecaa01c9075dc83ab8da95cfda3fb

Observation 1cbf2e62-bd01-40e5-8945-0f6bbbf8c937 · inbound

HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds cites this paper.

HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds SAM 3: Segment Anything with Concepts

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-10T13:45:27.276875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T13:45:24.961208Z digest=sha256:34e184bdeb8f8a50a913e5ac1ecaf18a737cba11b1d65913d25ea3bf1ff0dc94

Observation 0d52dcb8-965e-4f38-bb0c-c15c017c191c · inbound

Geometrically Consistent Multi-View Scene Generation from Freehand Sketches cites this paper.

Geometrically Consistent Multi-View Scene Generation from Freehand Sketches SAM 3: Segment Anything with Concepts

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T13:25:26.025915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T13:23:58.169493Z digest=sha256:55e4c0c3464592b36577b8f76ed08df0bafbfc2f484dc0271281c9ac1597396e

Observation 9c1b7aa7-1c15-487f-9810-a14ae8d8bd88 · inbound

Geometrically Consistent Multi-View Scene Generation from Freehand Sketches cites this paper.

Geometrically Consistent Multi-View Scene Generation from Freehand Sketches SAM 3: Segment Anything with Concepts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T16:18:25.431492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:18:25.431492Z digest=sha256:d7e0ce44ed3fd3e57f309207d887ce3cd7e9973f85b715f867f614c009d1d110

Observation 45d12334-47f7-402c-8d6c-0cf00521b1ee · inbound

HRDexDB: A Paired Human-Robot Dataset for Cross-Embodiment Dexterous Grasping cites this paper.

HRDexDB: A Paired Human-Robot Dataset for Cross-Embodiment Dexterous Grasping SAM 3: Segment Anything with Concepts

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-10T10:49:56.066534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T10:48:27.803794Z digest=sha256:e32bd7192c476a13c05d7793d512eeac6b3f17f8f4e9bf6ac5ba172bf410e578

Observation cccf5171-c325-4e0a-b3e4-5ef86068b754 · inbound

HRDexDB: A Paired Human-Robot Dataset for Cross-Embodiment Dexterous Grasping cites this paper.

HRDexDB: A Paired Human-Robot Dataset for Cross-Embodiment Dexterous Grasping SAM 3: Segment Anything with Concepts

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-12T19:56:39.725379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T19:56:39.725379Z digest=sha256:6bf8e51b539f340155473743b7cefbc972e2383105d9685bbd09a40ecbb8398a

Observation fb75c2a4-e3ae-4692-a011-dbe84b769df2 · inbound

AnimationBench: Are Video Models Good at Character-Centric Animation? cites this paper.

AnimationBench: Are Video Models Good at Character-Centric Animation? SAM 3: Segment Anything with Concepts

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T11:50:20.361212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T11:47:27.131584Z digest=sha256:310e2f3f8f30976440d02f74e24caab899a2d1917058b1f45dbb634651c1cd16

Observation 32a54687-4f8d-4393-aefe-39c01e1bff4b · inbound

Fringe Projection Based Vision Pipeline for Autonomous Hard Drive Disassembly cites this paper.

Fringe Projection Based Vision Pipeline for Autonomous Hard Drive Disassembly SAM 3: Segment Anything with Concepts

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-10T07:01:49.507373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:56:58.988003Z digest=sha256:a173e21c3adeb3b122a11f7b29448620562b2bad810c532bea047fe977ecef10

Observation f6859d33-06f2-4c06-95f5-47bb7e761261 · inbound

Real-Time Cellist Postural Evaluation With On-Device Computer Vision cites this paper.

Real-Time Cellist Postural Evaluation With On-Device Computer Vision SAM 3: Segment Anything with Concepts

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-10T05:41:02.090528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T05:38:35.025500Z digest=sha256:df83f1c1b2653ce9a5882079ac0f619ddd21e98e6607e1b31d73ae54ab4c70cf

Observation 4b0be7c8-f854-4344-a1cb-ebe14ff0cc55 · inbound

DiffuSAM: Diffusion Guided Zero-Shot Object Grounding for Remote Sensing Imagery cites this paper.

DiffuSAM: Diffusion Guided Zero-Shot Object Grounding for Remote Sensing Imagery SAM 3: Segment Anything with Concepts

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T10:09:08.371735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T05:06:06.159341Z digest=sha256:6cb81ef9464061cbe302d7a0dd1728a40375fc80545d151d6210c464fce6a7b3

Observation 362df0b7-cba3-4610-837c-a1504f26bbaf · inbound

Is SAM3 ready for pathology segmentation? cites this paper.

Is SAM3 ready for pathology segmentation? SAM 3: Segment Anything with Concepts

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-10T12:05:23.345458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T04:40:57.284414Z digest=sha256:99af79ca787556decf270829ea7e1f05d3786f0a3cf294c2141dd87d053cfda6

Observation 5cbea517-b97b-4613-bff1-5c6693c024e4 · inbound

Is SAM3 ready for pathology segmentation? cites this paper.

Is SAM3 ready for pathology segmentation? SAM 3: Segment Anything with Concepts

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:58:03.759666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T21:54:14.103479Z digest=sha256:acea990797bb56cd8ec50e21052e43b092564d40a5eb65f919bb83e3820f9d28

Observation b9a83476-dbcf-4bfc-8484-6763f09f7e8e · inbound

APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track cites this paper.

APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track SAM 3: Segment Anything with Concepts

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-10T03:29:21.560465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T03:29:12.783993Z digest=sha256:1a06f257c9f0316a7d913b8d1703cf7537465237258f5e8f7426604941d69d90

Observation 3c57f55c-6c02-44ac-91cb-7da2f25c5a26 · inbound

PLaMo 2.1-VL Technical Report cites this paper.

PLaMo 2.1-VL Technical Report SAM 3: Segment Anything with Concepts

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-11T12:46:02.530784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T03:09:05.436900Z digest=sha256:231e844f56717065e0b3f598e8f4d784ad5cada81414e8ca12a4e40130a83ab0

Observation 7da65e17-8820-4d66-b5b0-e1e53d27b9d8 · inbound

Structure-Semantic Decoupled Modulation of Global Geospatial Embeddings for High-Resolution Remote Sensing Mapping cites this paper.

Structure-Semantic Decoupled Modulation of Global Geospatial Embeddings for High-Resolution Remote Sensing Mapping SAM 3: Segment Anything with Concepts

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:16:06.030436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T02:06:05.390586Z digest=sha256:bf8ce6b031255880bee2d86f621d191c86554ef30d30b5df78a1e6d3a350825d

Observation ec16c773-e5db-44b5-a1a6-2b7cc9db6172 · inbound

Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding cites this paper.

Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding SAM 3: Segment Anything with Concepts

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-11T12:31:03.223382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T03:33:56.211514Z digest=sha256:8e0e6c2d39762fa11a432437f204f5402884231e772a88463beefd7c2f1be3a6

Observation 61af82f1-2a10-4679-8046-ee4ed96b151e · inbound

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation cites this paper.

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation SAM 3: Segment Anything with Concepts

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T12:41:04.398550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T03:14:45.834520Z digest=sha256:7911f86251af597432603fc441d8328265575da62f38a1a3755c72f3cab51bb5

Observation 0dbc6211-a1cc-4abe-9791-3e525f6677a2 · inbound

CoCo-SAM3: Harnessing Concept Conflict in Open-Vocabulary Semantic Segmentation cites this paper.

CoCo-SAM3: Harnessing Concept Conflict in Open-Vocabulary Semantic Segmentation SAM 3: Segment Anything with Concepts

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T03:14:08.261351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T03:10:10.398335Z digest=sha256:c713bf0fea44b736358833b590e3b2229a52b784c59ca6246556d0e8d7c3a09a

Observation 903d84b6-01e3-4107-b94d-d3c08964ba91 · inbound

Rethinking Where to Edit: Task-Aware Localization for Instruction-Based Image Editing cites this paper.

Rethinking Where to Edit: Task-Aware Localization for Instruction-Based Image Editing SAM 3: Segment Anything with Concepts

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-10T00:39:48.383774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T00:39:21.872643Z digest=sha256:e6c093e0f503c2876e60750476a7cddb5a815ec612aec000fda9c5730e55b892

Observation 4b49dd26-26a1-45f9-8555-f9c41b018145 · inbound

Image Generators are Generalist Vision Learners cites this paper.

Image Generators are Generalist Vision Learners SAM 3: Segment Anything with Concepts

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:41:06.000924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T01:14:05.034951Z digest=sha256:5936a9ea55a563245ac4df4eb9c883922e56ee4db2a8059768bb0318d71acbb7

Observation af3f4cfb-5577-4c1b-8d9a-e907ffa95274 · inbound

Image Generators are Generalist Vision Learners cites this paper.

Image Generators are Generalist Vision Learners SAM 3: Segment Anything with Concepts

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:45:14.665089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T07:40:46.090808Z digest=sha256:49025b2b2d7e459759b506d5ec818ea8216b5d352bfb77306f4b56d3092c3398

Observation a9475ea4-8305-4180-84e9-853d6caf6d10 · inbound

VFM$^{4}$SDG: Unveiling the Power of VFMs for Single-Domain Generalized Object Detection cites this paper.

VFM$^{4}$SDG: Unveiling the Power of VFMs for Single-Domain Generalized Object Detection SAM 3: Segment Anything with Concepts

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:31:07.242137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T21:38:15.877684Z digest=sha256:0fc5092ca14c6efc617961b1fd0a92a889a7245c2fe621cf9ac7dabed0487a83

Observation 7810396d-378d-4d90-a2ab-75b1d2f5e5dd · inbound

VFM$^{4}$SDG: Unveiling the Power of VFMs for Single-Domain Generalized Object Detection cites this paper.

VFM$^{4}$SDG: Unveiling the Power of VFMs for Single-Domain Generalized Object Detection SAM 3: Segment Anything with Concepts

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:05:26.336364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T06:03:56.040612Z digest=sha256:2d58b0a8698899bfe493687eceab16d36927c950971edc0bbf365947a8cf1f73

Observation d73ef529-973a-48c6-904c-7c4e31211a91 · inbound

Focus Session: Hardware and Software Techniques for Accelerating Multimodal Foundation Models cites this paper.

Focus Session: Hardware and Software Techniques for Accelerating Multimodal Foundation Models SAM 3: Segment Anything with Concepts

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-09T23:04:17.777657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T23:02:07.554154Z digest=sha256:23e500f9d1c34ff53d231fae9dfa20e1daf320ca0ac3b82c2f4b8c63d8cf8b06

Observation 10f282c6-061e-4efb-9c2a-0bcb6fcbebc3 · inbound

AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method cites this paper.

AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method SAM 3: Segment Anything with Concepts

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-10T09:33:41.849359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T05:14:25.423302Z digest=sha256:36c84213948ed24f44079bed80a935552b1cbd10b94b417a9fc741850877bc87

Observation f2f1cc56-c5dd-4618-a64e-a4a3eb0c0e6c · inbound

OAMVOS:2nd Report for 5th PVUW MOSE Track cites this paper.

OAMVOS:2nd Report for 5th PVUW MOSE Track SAM 3: Segment Anything with Concepts

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-10T10:09:08.582081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T05:04:56.366788Z digest=sha256:dcc9f0471dc4969eda1a1107fc0681eba69a14b1367a481553af0039b43e267f

Observation 43085a33-3f39-40ad-86d0-acc05b57c688 · inbound

SketchVLM: Vision language models can annotate images to explain thoughts and guide users cites this paper.

SketchVLM: Vision language models can annotate images to explain thoughts and guide users SAM 3: Segment Anything with Concepts

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T14:36:05.439172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T21:32:08.503584Z digest=sha256:b8dadba8ac62cda161d362039102b6499e4a252a9177894c68b95ede07ba7653

Observation 7dabd64c-9557-42c7-934f-77ebfa623672 · inbound

INSIGHT: Indoor Scene Intelligence from Geometric-Semantic Hierarchy Transfer for Public~Safety cites this paper.

INSIGHT: Indoor Scene Intelligence from Geometric-Semantic Hierarchy Transfer for Public~Safety SAM 3: Segment Anything with Concepts

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:31:14.589423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T08:38:36.288087Z digest=sha256:2e3b568328914dcb837e53ac450c681752bf7b281c75bea3db3387b737336ede

Observation 0fb5166b-55d3-42dc-bedb-745b02c5409e · inbound

BridgeACT: Bridging Human Demonstrations to Robot Actions via Unified Tool-Target Affordances cites this paper.

BridgeACT: Bridging Human Demonstrations to Robot Actions via Unified Tool-Target Affordances SAM 3: Segment Anything with Concepts

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:46:15.633882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T07:55:35.304883Z digest=sha256:66b8f2fd14ad919f410b3253238de33356965beb403d0f1072980ee556cf3de2