Pith. sign in

Paper Citation Record · LEDGER

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation

As of 7 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2507.05948.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.05948 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:17:34.548221Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy50
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 44cec9f3-2f55-4dd1-946f-2731467497a5 · outbound

This paper cites Stem-seg: Spatio-temporal em- beddings for instance segmentation in videos.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Stem-seg: Spatio-temporal em- beddings for instance segmentation in videos

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.225576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:30.801393Z digest=sha256:2812f00b453b8f3b8a554b43887a7ad11ff9e63f38832080141d24ea4d827028

Observation 4aa3b6fc-8dfd-4296-950a-de684d2ad91e · outbound

This paper cites Is space-time attention all you need for video understanding? In International Conference on Machine Learning, 2021.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Is space-time attention all you need for video understanding? In International Conference on Machine Learning, 2021

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.215551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:30.900956Z digest=sha256:ec1b8c2d0644a87c29231e971cadd44112d84c2e90d51ac5711a51f027abc774

Observation 4fd57a2b-e37a-4bd8-a81b-f11ca7fe49d5 · outbound

This paper cites MiDaS v3.1 -- A Model Zoo for Robust Monocular Relative Depth Estimation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation MiDaS v3.1 -- A Model Zoo for Robust Monocular Relative Depth Estimation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:17:31.044998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:17:31.044998Z digest=sha256:521cfdf74168561705bf354597eaf6c997034a06a3a2e328cac39a209dd33cdc

Observation f42aa1f6-bfb6-4184-b5cd-a5166450ac14 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:17:31.260099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:17:31.260099Z digest=sha256:974a83b4504a1f6f985fb124bda532d678a25f070afacab82f5b9add1f66f83d

Observation 40888369-7740-4196-8a40-0ca91b09051e · outbound

This paper cites End-to- end object detection with transformers.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation End-to- end object detection with transformers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.205943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:31.370901Z digest=sha256:03f1272094afc6438570ade5c5751902d00de936d57751afdaa0a7c1789c6416

Observation de672815-1933-4925-b478-5ca43aa28665 · outbound

This paper cites Liu, Yen-Cheng Liu, and Yu- Chiang Frank Wang.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Liu, Yen-Cheng Liu, and Yu- Chiang Frank Wang

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.195868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:31.495860Z digest=sha256:08295b5a77a08129b8cbcfef1428a731c3adf60480eba2f2d722dbb652f9c0bc

Observation 883b2de5-6e7e-4ce7-82b9-f04d83ff57dc · outbound

This paper cites Vision Transformer Adapter for Dense Predictions.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Vision Transformer Adapter for Dense Predictions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:17:31.571904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:17:31.571904Z digest=sha256:2ebe2cc8d8a112a57ad0974b56f637cd3023e80d5b6a2b0122e5e1d86b7a0a56

Observation bee4c805-d516-4b99-a360-31bdf9138c91 · outbound

This paper cites Collins, Yukun Zhu, Ting Liu, Thomas S.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Collins, Yukun Zhu, Ting Liu, Thomas S

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.185091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:31.648706Z digest=sha256:9fe3cf9e60f266b23fcd736c2162de0ee4499c5a8c759fb63f2562ba6ffb3786

Observation bd7d29b2-c6db-4d8c-b6fc-7cec9ca92856 · outbound

This paper cites Mask2Former for Video Instance Segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Mask2Former for Video Instance Segmentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:17:31.737828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:17:31.737828Z digest=sha256:71f6dc005244320fe7e4e69ce9fea4168e28c5d1b4a7ddba462148530dd6ec16

Observation f2b25763-c8f1-4016-af5e-0fa2158affd1 · outbound

This paper cites Per- pixel classification is not all you need for semantic segmen- tation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Per- pixel classification is not all you need for semantic segmen- tation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.174848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:31.802513Z digest=sha256:e83b0eb0d19d5d5f2daae1dba3abbb82f2ee1bd8c7c7e5f3b4748373e762d76c

Observation f6b9670c-78a4-47e4-9fc6-780bc041e9df · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Masked-attention mask transformer for universal image segmentation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.164477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:31.862573Z digest=sha256:a363a5518d3c8161213f352760ea864d588f2e607f86baf4b8c797bfa3ebaf2f

Observation eecf82d2-8b0c-41c0-bc50-4a1879bd9af6 · outbound

This paper cites MeViS: A large-scale benchmark for video segmentation with motion expressions.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation MeViS: A large-scale benchmark for video segmentation with motion expressions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.153426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:31.953003Z digest=sha256:ff676359cd6b2049f226ca545c1749179203b5e8b37dd8b849af08bce2d88fc8

Observation f81a655b-2269-4813-b6dd-62845d006d2c · outbound

This paper cites MOSE: A new dataset for video object segmentation in complex scenes.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation MOSE: A new dataset for video object segmentation in complex scenes

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.143479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:32.053673Z digest=sha256:60529a9d5c554b49df811dc549e39ed6b717978953348469a8e088b959b5abff

Observation 4ea2f2aa-c283-4462-aa95-184f0563ba7c · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation An image is worth 16x16 words: Transformers for image recognition at scale

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.132546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:32.120327Z digest=sha256:addd3bcaad4b2c0c12d16e5d7985591173232b19323d19a2dbfbaa001e1eb3c6

Observation 2037d9f0-e08e-4e28-a9a3-0c6ebf7a7c0c · outbound

This paper cites Depth map prediction from a single image using a multi-scale deep net- work.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Depth map prediction from a single image using a multi-scale deep net- work

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.121778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:32.186887Z digest=sha256:b6f2ea1a99b859d8d2d646f663795b0c3eaa84a163d6ff1ec2121b440f2fb0e0

Observation 19f79190-0953-41c4-9daa-0e0f6c04b649 · outbound

This paper cites Deep Ordinal Regression Network for Monocular Depth Estimation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Deep Ordinal Regression Network for Monocular Depth Estimation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.111325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:32.259930Z digest=sha256:e978c5faf90f8a3d65ffcebb1c98c597457d5186b0081d7beba1146b47538b06

Observation e1676818-0491-41a7-bc6c-4f4c2cc0a5ea · outbound

This paper cites Gratt-vis: Gated residual atten- tion for video instance segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Gratt-vis: Gated residual atten- tion for video instance segmentation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.100906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:32.350163Z digest=sha256:c1adf0cb5de5728931499e070c097c8b7a33d91a125e1179d13e5b7a6d294730

Observation dfe9f3e3-c457-4e73-a771-9c6023ce1132 · outbound

This paper cites Deep residual learning for image recognition.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Deep residual learning for image recognition

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.090612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:32.414701Z digest=sha256:47ea59e84ca08b5168a6de62f316dafbaff34a3c9b4ebbc07888ac9db1da697b

Observation b898921b-6756-44f9-b53b-4806e2b439ab · outbound

This paper cites Vita: Video instance segmentation via object token association.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Vita: Video instance segmentation via object token association

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.080147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:32.485519Z digest=sha256:156dbc2fbf67f568385701dbe8fa844843f75a989b6c6f350e358a739f5f5ac5

Observation d1b04145-c257-4b6e-8936-eb3f465a6564 · outbound

This paper cites A generalized framework for video instance segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation A generalized framework for video instance segmentation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.067855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:32.555234Z digest=sha256:e91e2cb2203c30f8c7cf4570375bd45d45b44031452898194f8b48008c3ea71a

Observation 47c16f76-4c88-48f2-8444-e17c89f1fdba · outbound

This paper cites Metric3d v2: A versatile monocular geomet- ric foundation model for zero-shot metric depth and surface normal estimation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Metric3d v2: A versatile monocular geomet- ric foundation model for zero-shot metric depth and surface normal estimation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.056740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:32.638086Z digest=sha256:9691792b3abb53b0b827353b06beaf3ee9173daae98186b83817b2e9dd0c2410

Observation a80c305d-816f-4cb4-92ea-bdd5366268d3 · outbound

This paper cites Min- vis: A minimal video instance segmentation framework without video-based training.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Min- vis: A minimal video instance segmentation framework without video-based training

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.046211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:32.704845Z digest=sha256:69f95fab3fdf022a3fa7617c00f1f6ef3cafde6050ff166a90df03d858cae79e

Observation 92e1425d-cafc-45f8-89f8-662caf0af0ba · outbound

This paper cites Video object segmentation with language referring expressions.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Video object segmentation with language referring expressions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.035711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:32.777215Z digest=sha256:d993e8fdd1963e88cf7397485657e0be3ca45b4dd09a0f5f3f3b2ae68128ce1a

Observation f48d6ef3-5bee-4933-8c66-12578057479c · outbound

This paper cites Self-Supervised Monocular Depth Es- timation: Solving the Dynamic Object Problem by Seman- tic Guidance.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Self-Supervised Monocular Depth Es- timation: Solving the Dynamic Object Problem by Seman- tic Guidance

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.023162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:32.849148Z digest=sha256:8c8ca7b57c01e40ba0736400b79bae4c7acd3646312759639544a2d848b4a5f8

Observation a105ac78-8cbf-4cd0-b8fd-1e2449f60562 · outbound

This paper cites CAVIS: Context-Aware Video Instance Segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation CAVIS: Context-Aware Video Instance Segmentation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:17:32.925664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:17:32.925664Z digest=sha256:e1bb61e8fa6c754b6bc477e519d3d1930010095750578a57c5cc2717a00379f1

Observation 3e00813c-e74e-4afb-85a3-4eeab842981a · outbound

This paper cites Tcovis: Temporally consistent online video instance seg- mentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Tcovis: Temporally consistent online video instance seg- mentation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.010380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:33.011814Z digest=sha256:67db118a427b1eef8a2242f665873fc11252e1611f1d21a639290dae474cdeb1

Observation 97a9959c-2aba-4451-9256-df2fac6f7eab · outbound

This paper cites Video k-net: A simple, strong, and unified baseline for video segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Video k-net: A simple, strong, and unified baseline for video segmentation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.998688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:33.092555Z digest=sha256:ca33b00ec675e12970642823644e7b7dbcf0681d3e2e78a263e9b8de435ff2ec

Observation 2f402b5d-1fa7-45e1-acea-a5cfffc5d4e9 · outbound

This paper cites Transformer-based visual segmenta- tion: A survey.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Transformer-based visual segmenta- tion: A survey

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.984574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:33.214189Z digest=sha256:d12e7d795562de256a0f0b077b128ad158662f5c8c22d084332f162586ac5b57

Observation 1fb97ea8-58a2-4419-a407-c80f8ccb521d · outbound

This paper cites Omg-seg: Is one model good enough for all segmentation? In Conference on Computer Vision and Pattern Recognition,.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Omg-seg: Is one model good enough for all segmentation? In Conference on Computer Vision and Pattern Recognition,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.972284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:33.279285Z digest=sha256:cfc8a1033ea575a32dcc97229e05f011fbff1bf1b9282a2a106e8465ea93e1c5

Observation 68151916-4bf7-463c-89ba-f2551fba9e48 · outbound

This paper cites Lawrence Zitnick.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Lawrence Zitnick

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:17:33.336373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:17:33.336373Z digest=sha256:179c071b9093c662772a2f7526988a7102bd3da015690f0684f52e05779a20db

Observation fb27e04d-cf28-4f62-9267-6313a358a888 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Swin transformer: Hierarchical vision transformer using shifted windows

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.956085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:33.406701Z digest=sha256:98b0a786d77531389364f0dbc154f9a9a6397bd3b96d139847cb1cb8299101ff

Observation 54b62558-b8a6-4d0a-938d-3233084b5af9 · outbound

This paper cites Decoupled weight de- cay regularization.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Decoupled weight de- cay regularization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:17:33.456111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:17:33.456111Z digest=sha256:3e3e97cd4e1ee7f465ed11a2074782dd1f02320757d65d3fd31d008f055422d7

Observation 8abc1bd3-3b74-4663-b440-c98b2466d8cb · outbound

This paper cites an unresolved cited work.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:17:34.937844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:33.525177Z digest=sha256:aeabf09a86b089b7bca5c081729e2070c126665708580212b7c1c3677ffb0d21

Observation b76ac901-acca-46a2-a6b2-c304a2344ab0 · outbound

This paper cites Keeping your eye on the ball: Trajec- tory attention in video transformers.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Keeping your eye on the ball: Trajec- tory attention in video transformers

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.926944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:33.620382Z digest=sha256:05a5d15e23cb782e52be67643b3c83ed4244bb60d2c4a36d787138e1ac25dc7c

Observation 4561aba0-e6f1-4af7-aea8-2c6015f83e70 · outbound

This paper cites A benchmark dataset and evaluation methodology for video object segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation A benchmark dataset and evaluation methodology for video object segmentation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.913748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:33.639897Z digest=sha256:8ece0fc3fab631a9cc6ff1e994b8ac6c7cb043a4b52466cb2d1e741f611aca08

Observation 98ad0c44-f3dc-4db5-96f9-c420c1c8def2 · outbound

This paper cites Occluded video instance segmentation: A bench- mark.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Occluded video instance segmentation: A bench- mark

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.902520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:33.718789Z digest=sha256:565fc9b548e7fae9a2802f7545881ca444c1c9630285e9863f24fd6c5cf6d60b

Observation 096fb90e-f66d-4bc0-aaff-31b9fe73e979 · outbound

This paper cites ViP-DeepLab: Learning Visual Perception with Depth-aware Video Panoptic Segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation ViP-DeepLab: Learning Visual Perception with Depth-aware Video Panoptic Segmentation

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:17:34.592459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:33.822660Z digest=sha256:88123c3f8fd9c7f5c7e20d95f29224e74adcada5ac77140650833bea6dae4881

Observation b8ff83cf-20d8-4991-94ad-7925dcb7cca5 · outbound

This paper cites Vi- sion transformers for dense prediction.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Vi- sion transformers for dense prediction

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.890407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:33.965005Z digest=sha256:4476fea1cd4869c59c92c46741c4708f79fa634d6d1001c67f8e107644b217ef

Observation 069c13b9-6382-42de-8880-d43d1eddbc4b · outbound

This paper cites Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.880084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:34.075884Z digest=sha256:63bd4bf96080e788dfb1c83f2f8bcaa3e215cfe04388c7dff05428380e88e105

Observation b6ed3fb2-ff4f-49d4-b10b-fec2342cbc41 · outbound

This paper cites Boosting monocular depth with panoptic segmentation maps.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Boosting monocular depth with panoptic segmentation maps

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.868877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:34.280261Z digest=sha256:ddc794f5d2906b747ad1282fc15100a3e9ae12646773f2e937fca5b6c16b515b

Observation faf64410-40a8-4a62-a17c-9a8c78694bee · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.856873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:34.460870Z digest=sha256:5c79689914a47425beaffc8df4dd9ff79addd753352f40978efa6aa8155dcb2a

Observation 2b611cf6-04e3-4099-900f-d70730844b99 · outbound

This paper cites Sigma: Siamese mamba network for multi-modal semantic segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Sigma: Siamese mamba network for multi-modal semantic segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.846739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:34.492979Z digest=sha256:2cb4c991954667fc1a1034428adc0e2ee5c85d7441f52c349620c7f27196c6b4

Observation f244ce5c-8d5f-48ac-8843-81dac7de53ce · outbound

This paper cites Sdc-depth: Semantic divide-and-conquer net- work for monocular depth estimation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Sdc-depth: Semantic divide-and-conquer net- work for monocular depth estimation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.835266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:34.496984Z digest=sha256:a70e8d53bd2dbbcaa00d86de3ab1b9b92b061b1eec81e535de84d21d27578960

Observation 85d59161-665f-4a1c-ac11-b17da801093c · outbound

This paper cites End-to-end video instance segmentation with transformers.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation End-to-end video instance segmentation with transformers

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.822760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:34.499751Z digest=sha256:dd1bbc30eabaf24adb1bc7a33c2e3fb693261c9a231ba80b818ade21cb08e837

Observation a4cb9d1e-976b-410b-9635-cc681dc350b4 · outbound

This paper cites Metric3d: Towards zero-shot metric 3d prediction from a single image.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Metric3d: Towards zero-shot metric 3d prediction from a single image

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.812977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:34.503601Z digest=sha256:36d2ffc493320e4df2a8c3529af895ce7c2c8ad7986a497ba63bcb95d0a5def7

Observation 44fad11c-bfcd-4554-bbf2-cabbca2bbdf3 · outbound

This paper cites Seqformer: Sequential transformer for video instance segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Seqformer: Sequential transformer for video instance segmentation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.803309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:34.507477Z digest=sha256:04aa7cae436c0b1e67aa46b1f8724daa30ef84929a09141484f7dd351187ecf8

Observation 0986d5df-f94b-4fe6-beb2-80d01711e7f5 · outbound

This paper cites In defense of online models for video instance segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation In defense of online models for video instance segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.792784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:34.510411Z digest=sha256:e55be16fdc930c2ffc2392d374d743cc551f1a5b3e67410f404a955f635a1d6f

Observation a4435c75-3b7d-47ec-a1ab-a37ebf32d005 · outbound

This paper cites Depth Any Video with Scalable Synthetic Data.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Depth Any Video with Scalable Synthetic Data

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:17:34.513382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:17:34.513382Z digest=sha256:2018292f313e2eb390a27bb2f42854525230846a8df395d3c6b15d36efdfe8b8

Observation b603f057-5f35-41a2-9a6e-ca11f6a59264 · outbound

This paper cites Video instance seg- mentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Video instance seg- mentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.781260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:34.516735Z digest=sha256:7c4d632e4f670d34b2d35fcb709bb95e4b56886642aa65b06ad7dc24a8189b3e

Observation 5ce503cb-9725-4c3e-bc4e-48fdd850bb6d · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Depth anything: Unleashing the power of large-scale unlabeled data

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.769290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:34.519533Z digest=sha256:ff4e125efcd8857d899f60afe1508f30b6b5ac9552908068f6e877bbff2a1957

Observation 249a832b-6022-41b9-bfe0-e82b34dad44e · outbound

This paper cites Depth any- thing v2.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Depth any- thing v2

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.757631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:34.522356Z digest=sha256:13428393e0c22a5a3050c2652424b98d15275b40bedd88dc9eea578e88462169

Observation 43c01758-bda8-406c-9a97-f7f7c2c6ee56 · outbound

This paper cites Ctvis: Consistent train- ing for online video instance segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Ctvis: Consistent train- ing for online video instance segmentation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.745864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:34.525132Z digest=sha256:a9d499da979c35a724e2fc82ae0e45ee452cf12bf2594209c1867c92b8ee2194

Observation 9448895f-a749-442b-907e-7e7bed8179af · outbound

This paper cites Polyphonicformer: Unified query learning for depth-aware video panoptic segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Polyphonicformer: Unified query learning for depth-aware video panoptic segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.733631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:34.528509Z digest=sha256:3106bf0b88a5adac0047cf74d981b287c071352af3b27d826113b5f455423fc8

Observation ee2c0955-4265-45f8-bb61-2b8b621a1eae · outbound

This paper cites Geometry meets semantic for semi-supervised monocular depth estimation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Geometry meets semantic for semi-supervised monocular depth estimation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.721040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:34.531650Z digest=sha256:6825c318a651f654974a4ad8f5c5c6ba437d33ca4abc2c26a4b382eedebca51b

Observation cf7c2cbb-991d-4625-a445-749a2e551d55 · outbound

This paper cites Cmx: Cross-modal fusion for rgb-x semantic segmentation with transformers.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Cmx: Cross-modal fusion for rgb-x semantic segmentation with transformers

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.710318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:34.534259Z digest=sha256:a8b26ff297986a4937326c525151f6f6980bed1be46d8808d113381cc07c49ac

Observation 33a3296b-c367-409b-b49a-db6338d5efd5 · outbound

This paper cites Delivering arbitrary-modal semantic segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Delivering arbitrary-modal semantic segmentation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.698098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:34.536874Z digest=sha256:8cdd60096dd55d208e8f456393968e3d7466ed9583330e53713097714d0e89c2

Observation 4ba3d615-4146-414e-975e-0459af2efcc2 · outbound

This paper cites Dvis: Decoupled video in- stance segmentation framework.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Dvis: Decoupled video in- stance segmentation framework

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.688108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:34.539546Z digest=sha256:a8166f2c68367c179f7e6972b159a5273c0a8ec2fa06539b02e24ac50f2a081b

Observation 648c8882-4c46-458c-8834-57f0771046d8 · outbound

This paper cites Dvis++: Improved decoupled frame- work for universal video segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Dvis++: Improved decoupled frame- work for universal video segmentation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.677425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:34.542429Z digest=sha256:20e4ba56d2a1eb7812d3f6a3276ce79d1d8d0e6afd00765cbdfc66588708c265

Observation fc0301be-19a4-4b39-a3f9-c4686cf35dbe · outbound

This paper cites Dvis-daq: Improving video segmentation via dynamic anchor queries.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Dvis-daq: Improving video segmentation via dynamic anchor queries

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.666323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:34.545623Z digest=sha256:843894836494d83a2cd0293308ed39c88406e1b5757680d5560243c29fd2e43b

Observation f903c52f-5a6a-4f45-9b73-c3fdb9922883 · outbound

This paper cites Deformable detr: Deformable transformers for end-to-end object detection.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Deformable detr: Deformable transformers for end-to-end object detection

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.655837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:17:34.548221Z digest=sha256:c66f3acd0937a7ce73a1f73727fb48f2a6b1947bbd0d963d2aafa11214e0aa6e

Pith citing papers

No inbound Pith citation observations are available.