Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-09T20:47:06.698475Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2605.00891.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-09T20:47:06.698475Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 189a533b-936a-4955-8d15-8da0e87348e6 · outbound
X2SAM: Any Segmentation in Images and Videos Qwen Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ee46ec15-f3da-40f0-beb6-71859a1d303e · outbound
X2SAM: Any Segmentation in Images and Videos LLaMA: Open and Efficient Foundation Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 112e7e39-cdcb-4a74-bed1-3d0328ae86cd · outbound
X2SAM: Any Segmentation in Images and Videos Learning transferable visual models from natural language supervision
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d661328e-2fc3-4b90-9f66-18185220cd93 · outbound
X2SAM: Any Segmentation in Images and Videos Scaling up visual and vision-language representation learning with noisy text supervision
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1a9c4b47-0ff5-4ef6-9da4-ce5691db9d97 · outbound
X2SAM: Any Segmentation in Images and Videos Show, attend and tell: Neural image caption generation with visual attention
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation efbb2c68-ac56-46c8-9834-5c011cd288bd · outbound
X2SAM: Any Segmentation in Images and Videos Vqa: Visual question answering
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8a666af3-d6f6-40ee-95c4-935b067aeaff · outbound
X2SAM: Any Segmentation in Images and Videos Language-based image editing with recurrent attentive models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 73d5b524-410c-4d1c-a816-5c329c2a46cb · outbound
X2SAM: Any Segmentation in Images and Videos Segment anything
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 57511c89-3be4-4704-b86d-a1b3c87186d1 · outbound
X2SAM: Any Segmentation in Images and Videos SAM 2: Segment Anything in Images and Videos
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9a1001aa-4c7b-42f2-8239-3d622cc1ab50 · outbound
X2SAM: Any Segmentation in Images and Videos Lisa: Reasoning segmentation via large language model
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 83a0d30e-093f-4cc7-8039-969fd5a5b955 · outbound
X2SAM: Any Segmentation in Images and Videos Visa: Reasoning video object segmentation via large language models.ECCV
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d02ec514-18dc-494d-abc9-69eba5c3511d · outbound
X2SAM: Any Segmentation in Images and Videos One token to seg them all: Language instructed reasoning segmentation in videos.NeurIPS
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ef49ba74-329f-4077-92a8-5222ed275bfb · outbound
X2SAM: Any Segmentation in Images and Videos Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 843c713f-225c-47ee-b7f7-0f8ab3275bbc · outbound
X2SAM: Any Segmentation in Images and Videos Improved baselines with visual instruction tuning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c95311b6-83b1-40e2-8254-c9c2a49a22b4 · outbound
X2SAM: Any Segmentation in Images and Videos Tarvis: A unified approach for target-based video segmentation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aeba86b3-ffcb-4e68-b3b7-56a9ca1126c1 · outbound
X2SAM: Any Segmentation in Images and Videos Oneformer: One transformer to rule universal image segmentation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2df7d21c-1b2b-4b3b-9f7c-1ade3bd75d29 · outbound
X2SAM: Any Segmentation in Images and Videos Omg-seg: Is one model good enough for all segmentation? InCVPR, pages 27948–27959
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c7c79173-c57b-4342-bb2f-8493a602bb72 · outbound
X2SAM: Any Segmentation in Images and Videos Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding.NeurIPS, 37:71737–71767
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 00024a58-bc61-40ad-8391-b81693f6e39e · outbound
X2SAM: Any Segmentation in Images and Videos Temporal memory attention for video semantic segmentation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f512585c-0589-43c6-ae74-50191e885586 · outbound
X2SAM: Any Segmentation in Images and Videos Video k-net: A simple, strong, and unified baseline for video segmentation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation af656a47-b134-4013-b0d3-5e5c5ed71abe · outbound
X2SAM: Any Segmentation in Images and Videos X-sam: From segment anything to any segmentation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6974a5d1-5310-4dd2-baeb-57ed19f590ea · outbound
X2SAM: Any Segmentation in Images and Videos Visual instruction tuning.NeurIPS, 36:34892–34916
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 64660bb0-4f1c-4366-924a-f0b000bbab47 · outbound
X2SAM: Any Segmentation in Images and Videos Llavanext: Improved reasoning, ocr, and world knowledge
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 02ce4c79-32d9-49c6-941c-5a1e0530e1d5 · outbound
X2SAM: Any Segmentation in Images and Videos LLaVA-OneVision: Easy Visual Task Transfer
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c1ff8307-044d-4a7e-be60-6a2c5bfcf9cf · outbound
X2SAM: Any Segmentation in Images and Videos How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 896b17eb-d5e7-4d8f-939c-172841dd903a · outbound
X2SAM: Any Segmentation in Images and Videos Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8a919a35-e3c3-4889-8b0a-33b4989eb4b0 · outbound
X2SAM: Any Segmentation in Images and Videos Glamm: Pixel grounding large multimodal model
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0196f2a1-bdb6-4125-8449-43c5cd76b280 · outbound
X2SAM: Any Segmentation in Images and Videos Videoglamm: A large multimodal model for pixel-level visual grounding in videos.CVPR
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation af0980ee-7629-4667-9e72-759526988977 · outbound
X2SAM: Any Segmentation in Images and Videos Psalm: Pixelwise segmentation with large multi-modal model
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bba4b529-0457-41d8-84ff-7f7bc8d4facf · outbound
X2SAM: Any Segmentation in Images and Videos HyperSeg: Towards Universal Visual Segmentation with Large Language Model
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b81759f9-1147-4d4a-819b-121ec2ec0a0b · outbound
X2SAM: Any Segmentation in Images and Videos Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f2b0c637-ef50-418c-992d-c7bd6da00cb1 · outbound
X2SAM: Any Segmentation in Images and Videos Qwen3-VL Technical Report
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ec346914-1640-4eef-afb5-371ccabdbf0e · outbound
X2SAM: Any Segmentation in Images and Videos Ferret: Refer and Ground Anything Anywhere at Any Granularity
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 426e0b0d-f9b4-4e99-9fbf-3abad408de99 · outbound
X2SAM: Any Segmentation in Images and Videos V-net: Fully convolutional neural networks for volumetric medical image segmentation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0e9620a1-d871-43b7-bbab-e92b044516c7 · outbound
X2SAM: Any Segmentation in Images and Videos Improving language understanding by generative pre-training
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 48ae41de-27a5-4684-b2ce-c0d10650fb2b · outbound
X2SAM: Any Segmentation in Images and Videos Focal loss for dense object detection
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5714d776-dbe0-48c9-a429-e69de5135a60 · outbound
X2SAM: Any Segmentation in Images and Videos Large-scale video panoptic segmentation in the wild: A benchmark
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ecd7e4e6-cd0b-47d2-be74-ce892fff8d39 · outbound
X2SAM: Any Segmentation in Images and Videos Vspw: A large-scale dataset for video scene parsing in the wild
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b59f63e8-6e66-4f52-ba88-c9b6cacba9d5 · outbound
X2SAM: Any Segmentation in Images and Videos Video instance segmentation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9cf3bdca-9cc3-4380-b93c-f4fa1c233029 · outbound
X2SAM: Any Segmentation in Images and Videos Urvos: Unified referring video object segmentation network with a large-scale benchmark
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2cb0fd27-26fd-4a0c-a948-00e046bcd75b · outbound
X2SAM: Any Segmentation in Images and Videos YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1b3a79e1-793e-47b5-bfbb-8acd67dd255f · outbound
X2SAM: Any Segmentation in Images and Videos A benchmark dataset and evaluation methodology for video object segmentation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 979af393-7e8b-4cc5-8da4-0f631655a1df · outbound
X2SAM: Any Segmentation in Images and Videos Video-chatgpt: Towards detailed video understanding via large vision and language models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4ff49862-0d0c-4d00-bc18-6d5cc73c35f3 · outbound
X2SAM: Any Segmentation in Images and Videos Gres: Generalized referring expression segmentation
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c65917f6-a72f-48a0-9a66-37df14f5e1ea · outbound
X2SAM: Any Segmentation in Images and Videos Semantic understanding of scenes through the ade20k dataset.IJCV, 127(3):302–321
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a8702369-aeaf-4897-814e-cf6494515560 · outbound
X2SAM: Any Segmentation in Images and Videos LoRA: Low-rank adaptation of large language models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d1bf84a3-8b50-4e07-bb31-8e73d1a805c3 · outbound
X2SAM: Any Segmentation in Images and Videos Decoupled Weight Decay Regularization
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 185942e1-e13f-4040-8fc0-7373101d478f · outbound
X2SAM: Any Segmentation in Images and Videos Open-vocabulary panoptic segmentation with text-to-image diffusion models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5775c423-951a-4a6d-866a-f7911a22f73e · outbound
X2SAM: Any Segmentation in Images and Videos Uniref++: Segment every reference object in spatial and temporal spaces.ICCV
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ddd436b2-739a-471c-bb30-79590767642d · outbound
X2SAM: Any Segmentation in Images and Videos Unipixel: Unified object referring and segmentation for pixel-level visual reasoning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5442f6fb-7a08-4af4-be2b-a5640366d55e · outbound
X2SAM: Any Segmentation in Images and Videos Segment everything everywhere all at once.NeurIPS, 36:19769–19782
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9bc49f70-74a6-4281-ae04-8ba5b8871e34 · outbound
X2SAM: Any Segmentation in Images and Videos Language as queries for referring video object segmentation
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2af7add8-b3e8-4b47-9326-6b86450a8b53 · outbound
X2SAM: Any Segmentation in Images and Videos LaSagnA: Language-based Segmentation Assistant for Complex Queries
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fc734836-8251-4641-914b-d10e09e118b0 · outbound
X2SAM: Any Segmentation in Images and Videos Mmbench: Is your multi-modal model an all-around player? InECCV, pages 216–233
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f03881bb-577e-44d7-8196-ef4ba1862f66 · outbound
X2SAM: Any Segmentation in Images and Videos Seed-bench: Benchmarking multimodal large language models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5fdedad0-1ec3-4043-96e2-c902833b1ce8 · outbound
X2SAM: Any Segmentation in Images and Videos Evaluating Object Hallucination in Large Vision-Language Models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9068ae36-3707-43b3-ac8a-39d266f118ab · outbound
X2SAM: Any Segmentation in Images and Videos A diagram is worth a dozen images
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d7dc8b0b-ed25-4f49-90b6-46dbdaae1da6 · outbound
X2SAM: Any Segmentation in Images and Videos Microsoft coco: Common objects in context
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 01815b34-207c-46b7-a5d7-16668649f050 · outbound
X2SAM: Any Segmentation in Images and Videos Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 671e29d7-b81b-4d47-b747-e2a39e9f0744 · outbound
X2SAM: Any Segmentation in Images and Videos Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9f13f2ab-6e5e-450e-92a2-663be295ff92 · outbound
X2SAM: Any Segmentation in Images and Videos Mlvu: Benchmarking multi-task long video understanding
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 85d37fd3-c47c-45a1-8d4b-29d6ce4dc423 · outbound
X2SAM: Any Segmentation in Images and Videos Longvideobench: A benchmark for long-context interleaved video-language understanding.NeurIPS, 37:28828–28857
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 03bffbe7-5c54-49a6-8635-8d9a376f551f · outbound
X2SAM: Any Segmentation in Images and Videos Masked-attention mask transformer for universal image segmentation
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0d3a54b7-25d4-4926-8063-01edbde880b2 · outbound
X2SAM: Any Segmentation in Images and Videos Cris: Clip-driven referring image segmentation
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 94abc434-0dfc-4bc3-ace5-c915712e418f · outbound
X2SAM: Any Segmentation in Images and Videos PG-Video-LLaVA: Pixel Grounding Large Video-Language Models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a2142cf6-256c-4bea-b154-a585cc137bd9 · outbound
X2SAM: Any Segmentation in Images and Videos Videochat: Chat-centric video understanding.Science China Information Sciences, 68(10):200102
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7feb229b-18d6-4698-89c7-3d85fe25ea84 · outbound
X2SAM: Any Segmentation in Images and Videos Chat-univi: Unified visual representation empowers large language models with image and video understanding
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 167756bb-13c3-4fc1-b365-064c3411a38c · outbound
X2SAM: Any Segmentation in Images and Videos Vila: On pre-training for visual language models
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
No inbound Pith citation observations are available.