Pith. sign in

Paper Citation Record · LEDGER

Native Segmentation Vision Transformers

As of 8 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 1 inbound Pith citation observation for arXiv:2505.16993.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16993 v1

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:55:40.776768Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T06:02:40.158866Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T06:07:22.539904Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact0
  • verified fuzzy53
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2460445c-9575-432c-a64a-35e294bcf3ca · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Native Segmentation Vision Transformers Swin transformer: Hierarchical vision transformer using shifted windows

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:34.491663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:34.491663Z digest=sha256:6e4f711715f294e0ea74d53061f1897e2a525ce545db718f55b6d9cb2fba4694

Observation 39342ace-32b9-4815-9690-66756b5c6ab3 · outbound

This paper cites Convnext v2: Co-designing and scaling convnets with masked autoencoders.

Native Segmentation Vision Transformers Convnext v2: Co-designing and scaling convnets with masked autoencoders

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:50.558282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:34.575933Z digest=sha256:adfe09b4cd6b8fbc0871e95ecb7f723dc6af5ef00c4abd24b928160f3f65adf4

Observation f518c955-48b5-4dc5-a695-53331bd64673 · outbound

This paper cites Neighborhood attention transformer.

Native Segmentation Vision Transformers Neighborhood attention transformer

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:50.429716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:34.685306Z digest=sha256:7baa3758c95c6654600c9b9f3ba47f00fee67f877be14f3a083bca500e9d6842

Observation 969585b4-5aa4-4515-985d-81316a64135c · outbound

This paper cites Backpropagation applied to handwritten zip code recognition.

Native Segmentation Vision Transformers Backpropagation applied to handwritten zip code recognition

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:34.815187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:34.815187Z digest=sha256:d585fa93a0a1ede2982b2ea2f339e9932d1d10f1a55694a9bb3903f9bb207e6a

Observation a9843379-0ad2-44e0-a301-20e7a7d3e85d · outbound

This paper cites an unresolved cited work.

Native Segmentation Vision Transformers Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:34.922357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:34.922357Z digest=sha256:b0cb85f69290847317f59b5f5f90e76bb87fb248d0aecdc58b7ca98f1b71e2e0

Observation da6f51d4-8ab3-42fd-99f6-aed459b366ac · outbound

This paper cites Schwing, Alexander Kirillov, and Rohit Girdhar.

Native Segmentation Vision Transformers Schwing, Alexander Kirillov, and Rohit Girdhar

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:35.010276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:35.010276Z digest=sha256:cda8b0108b17b76eeead7b4b1307d290fe5eadb437b52fa29b6cc809fe33d6a2

Observation cdb95e00-622d-4ef2-8539-659d45743bd9 · outbound

This paper cites Feature pyramid networks for object detection.

Native Segmentation Vision Transformers Feature pyramid networks for object detection

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:35.085000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:35.085000Z digest=sha256:7e5c37966b8f48427450b5fa8f7dbd6d234f3452e848e7fd7fdcffa698e6df64

Observation 991767dd-4efc-4e69-9305-9957da71bd7f · outbound

This paper cites FaPN: Feature-aligned pyramid network for dense image prediction.

Native Segmentation Vision Transformers FaPN: Feature-aligned pyramid network for dense image prediction

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:50.262110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:35.224616Z digest=sha256:3e06b47b59d49e86a80cc1911c4d66f665a8e21c200a125bfbd9a2cafd09ee26

Observation ee282015-9ab6-4dc9-a550-1d84fc6a92af · outbound

This paper cites an unresolved cited work.

Native Segmentation Vision Transformers Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:55:50.132482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:35.321961Z digest=sha256:6215c586ee943d432f28ff1f7c23e0359c31b4074b242244e224d3a4eb5e7800

Observation 3e5f1493-744d-47ed-934f-63537f47f13b · outbound

This paper cites Clusterformer: Clustering as a universal visual learner.

Native Segmentation Vision Transformers Clusterformer: Clustering as a universal visual learner

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:49.967714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:35.479711Z digest=sha256:905a68d0595bad8a6027b1d60c90df2e0f7fa382ed3b2dafb518bfad4190cf50

Observation fa701606-54c5-4881-b075-b7c7e1865346 · outbound

This paper cites Learning hierarchical image segmentation for recognition and by recognition.

Native Segmentation Vision Transformers Learning hierarchical image segmentation for recognition and by recognition

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:49.870902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:35.562953Z digest=sha256:0ab6a02e584b45db48fc26793d4a08b8ab1cc0d81f98c3d409fd5e5cedc24067

Observation 15597a3d-7613-4f33-8943-a46cfd1b6687 · outbound

This paper cites Tcformer: Visual recognition via token clustering transformer.

Native Segmentation Vision Transformers Tcformer: Visual recognition via token clustering transformer

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:49.742996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:35.643011Z digest=sha256:5566497f446881f43f0b4aa70b919252ab18a6cf4b63e8da36f36d10fca64351

Observation 2d765e73-8eea-41de-9bc5-055e8bb36032 · outbound

This paper cites Schwing, and Alexander Kirillov.

Native Segmentation Vision Transformers Schwing, and Alexander Kirillov

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:49.604488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:35.735251Z digest=sha256:2a237e669a5f019da31e9780299746b0fc274ea7aa0e26ab2bc1d84b185244e8

Observation 6f38e803-886f-4711-aeef-f7b547b01593 · outbound

This paper cites Slic superpixels compared to state-of-the-art superpixel methods.

Native Segmentation Vision Transformers Slic superpixels compared to state-of-the-art superpixel methods

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:49.402810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:35.837086Z digest=sha256:d14802971d39d89cd98f50a9599236e5466457352ec9fddaeb013c9020db9412

Observation 8bd76a18-d8a2-42f5-999d-27a3bca78879 · outbound

This paper cites Object-centric learning with slot attention.

Native Segmentation Vision Transformers Object-centric learning with slot attention

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:49.233922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:35.935254Z digest=sha256:dd462bf1f9187aef3ffb60af904350a9fce0cf017515e10a6f7323fca18cae13

Observation 2fa49496-b243-4745-a451-4f5bb7571d4b · outbound

This paper cites Mean shift: A robust approach toward feature space analysis.

Native Segmentation Vision Transformers Mean shift: A robust approach toward feature space analysis

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:49.016911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:36.074370Z digest=sha256:b621bf7074f9e3f7621743159f6ea4285359efc53c23bbc3ee7707c313a40804

Observation f1abbd73-3554-4b10-82bc-659c6ddb72ae · outbound

This paper cites Felzenszwalb and Daniel P.

Native Segmentation Vision Transformers Felzenszwalb and Daniel P

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:48.853816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:36.168668Z digest=sha256:e49e389e3c11e9c48fdd2126f1cbbdcdf33e2201a45753dee7b3a4bce9e39dcc

Observation b6cc0892-2632-4d79-a46b-dea5645c87d8 · outbound

This paper cites Contour detection and hierarchical image segmentation.

Native Segmentation Vision Transformers Contour detection and hierarchical image segmentation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:48.710783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:36.279399Z digest=sha256:cade63555aef6085dcd233c3f63ba3a0cbe78ff110aa44b05c5dce3b08408db6

Observation 6909a4cd-4635-4a9c-b02d-8368068fba64 · outbound

This paper cites Seeds: Superpixels extracted via energy-driven sampling.

Native Segmentation Vision Transformers Seeds: Superpixels extracted via energy-driven sampling

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:48.536144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:36.403076Z digest=sha256:1f62025aaa339dab4ad5b9db4db91fedbc1a8fcf2fe8bd688c44f5416560065d

Observation 6f8390d9-e7de-427f-ba63-7e7c1684f633 · outbound

This paper cites Semantic understanding of scenes through the ade20k dataset.

Native Segmentation Vision Transformers Semantic understanding of scenes through the ade20k dataset

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:48.395546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:36.473035Z digest=sha256:77c14960adffb8f43aa0a36d5d266dc0b434ff3d7220849b093d177e0d92ebdf

Observation 22559590-2a4f-410d-bdfa-744dd9271f2f · outbound

This paper cites Panoptic segmentation.

Native Segmentation Vision Transformers Panoptic segmentation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:48.202978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:36.569923Z digest=sha256:6a723fd26c2cb5da6dce13b362cb446f92f49beaced327a9759cca5321359298

Observation cc44dad6-1121-4f2f-a77c-7b88b2bded88 · outbound

This paper cites Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position.

Native Segmentation Vision Transformers Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:36.693822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:36.693822Z digest=sha256:f6acc692e303905750ac825b569862f2062e2e22a5860f769be1e50550a3c6a3

Observation 60d9cf58-f5f8-4f77-9266-976677bb21c0 · outbound

This paper cites Gradient-based learning applied to document recognition.

Native Segmentation Vision Transformers Gradient-based learning applied to document recognition

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:48.038344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:36.746616Z digest=sha256:2c971b598ab60082990a2e5dbe842185415e12e720d230aaadc8a9eb4d662027

Observation 11364c9b-1f78-438f-add2-63c1c7732b99 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Native Segmentation Vision Transformers An image is worth 16x16 words: Transformers for image recognition at scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:36.825137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:36.825137Z digest=sha256:ae3d1dff0db4ad77f6c1151255ac3a8bdda8dc8fa662ca2dd2d73cf8d12d0203

Observation edb1c32d-38b4-47cb-b971-d50f5cdb5fe7 · outbound

This paper cites A convnet for the 2020s.

Native Segmentation Vision Transformers A convnet for the 2020s

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:36.922349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:36.922349Z digest=sha256:8dfdfafb74f5bd15595ed33b39af13fb51f77396db201774ae53534a32608941

Observation 0350cb82-c86e-4800-bd90-f7e8a3f9b611 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Native Segmentation Vision Transformers SAM 2: Segment Anything in Images and Videos

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:37.004068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:37.004068Z digest=sha256:9cd37635eaf55cb3fd5dcfeaee3c5c33725f9d0cb0ba0ca4f8ab7f515cc8b346

Observation 9851423e-6728-49b0-a99d-6c901950ccb1 · outbound

This paper cites Fully convolutional networks for semantic segmentation.

Native Segmentation Vision Transformers Fully convolutional networks for semantic segmentation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:37.099662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:37.099662Z digest=sha256:8cf1bbea2144ccfde429083a6aa09280e75cc50bc60443f90a0a8047f7db8071

Observation 8e70d7bb-005f-436d-adb4-ffaa0f5b9655 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

Native Segmentation Vision Transformers U-net: Convolutional networks for biomedical image segmentation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:47.850044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:37.172521Z digest=sha256:ee0572ebf9e3aef2ea684a3cee6f6ff49cfedbf2547b4187af11ed844f9041d5

Observation 7296edd4-bfdd-41fa-b45b-d1ef3d797e07 · outbound

This paper cites Rich feature hierarchies for accurate object detection and semantic segmentation.

Native Segmentation Vision Transformers Rich feature hierarchies for accurate object detection and semantic segmentation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:47.634897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:37.259829Z digest=sha256:876362bd7158f6f9736a4ed81b68d94f50945ea0b43f2a27c313519cba853e1b

Observation 4617a109-394d-4a8d-b17d-2a6863660fa7 · outbound

This paper cites Fast r-cnn.

Native Segmentation Vision Transformers Fast r-cnn

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:37.360904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:37.360904Z digest=sha256:24028c2c9aa1fd29e5d8eaabeaca9ef40973ba136a19418772b9161f2805415b

Observation abbb1278-5402-40c8-b8a9-a796dda25b68 · outbound

This paper cites End-to-end object detection with transformers.

Native Segmentation Vision Transformers End-to-end object detection with transformers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:37.466195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:37.466195Z digest=sha256:9722bbb73c1450b200f080964e0f994d42655c0946ba05cd8fc2e91b2f2f9798

Observation e95186b4-c75a-4946-8c30-eb2aae10353d · outbound

This paper cites Normalized cuts and image segmentation.

Native Segmentation Vision Transformers Normalized cuts and image segmentation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:47.416823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:37.548532Z digest=sha256:8eeea2d8ecec4c9e58c84fc53274d5e688db43447fd885a88fa6f34ddcc3a24e

Observation 2a1bbe62-2422-4b2f-9320-d17f37cbd3cb · outbound

This paper cites Multiscale combinatorial grouping.

Native Segmentation Vision Transformers Multiscale combinatorial grouping

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:47.198109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:37.643606Z digest=sha256:e6e5e7c37d3bfc506b5f0e81aa887b2cb5ff0807d32e14b731e24115a743199b

Observation 46dc9746-5865-474f-b668-83260ca4fe38 · outbound

This paper cites Superpixel samping networks.

Native Segmentation Vision Transformers Superpixel samping networks

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:47.016623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:37.702010Z digest=sha256:3948ae92a74d43ac4103a00bf34c02e628794ac3fb41e7aa2545c38f49c41429

Observation abda8ddf-d305-4097-8a72-bf103965d096 · outbound

This paper cites an unresolved cited work.

Native Segmentation Vision Transformers Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:55:46.802461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:37.819714Z digest=sha256:1faee895b4d63ee2768f3b7e196ee8fa07e5f33ba8220d04ff92bcdb7dd4bdac

Observation db5cc85c-a4f2-48a8-b46e-064f879dc9a9 · outbound

This paper cites an unresolved cited work.

Native Segmentation Vision Transformers Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:55:46.557896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:37.930389Z digest=sha256:29ac8ff69ee368d7fadc75f889a1e7d9babc0eee1618d69c021b480c2581ccf1

Observation 9eca3a2d-053d-4021-b61b-b205ef653898 · outbound

This paper cites Unified perceptual parsing for scene understanding.

Native Segmentation Vision Transformers Unified perceptual parsing for scene understanding

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:46.374569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:38.018959Z digest=sha256:e61d6b554933af7329e1225b1215c6d11a3cbed21cdb0b944bfee10c55329d98

Observation 58de59c6-c80f-492b-a683-8e12dfa77e14 · outbound

This paper cites Berg, and Li Fei-Fei.

Native Segmentation Vision Transformers Berg, and Li Fei-Fei

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:46.201278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:38.112464Z digest=sha256:4a8e3b5385fbc7b45e9e56cb9f9cc51cc22dc6d467b5088e7814a398b38a2c1b

Observation e4bb2b64-dfeb-44b8-8a9f-32b9a46018a4 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Native Segmentation Vision Transformers Learning transferable visual models from natural language supervision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:38.182246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:38.182246Z digest=sha256:eb0b20e6847959e008af8063a2a9f27cb804ef00809f61ba1f67958e59846637

Observation 051e48ba-726e-45b9-b577-5958b6803ef4 · outbound

This paper cites Le, Yun- Hsuan Sung, Zhen Li, and Tom Duerig.

Native Segmentation Vision Transformers Le, Yun- Hsuan Sung, Zhen Li, and Tom Duerig

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:46.027620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:38.254570Z digest=sha256:8df4653fc345f6af364a7dd8addb6d638b822d1b5f9b138299b9d5225e3905e3

Observation 4c5dfd3e-7978-414e-93c7-3a0c2cfde865 · outbound

This paper cites A simple framework for text-supervised semantic segmentation.

Native Segmentation Vision Transformers A simple framework for text-supervised semantic segmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:45.829111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:38.330391Z digest=sha256:74e3d59072cf757f2d71939a5a1f76702d1245bb22feec7cf688823a53b3b102

Observation 522dc38d-724f-4e3d-bd7d-5badc2072f75 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

Native Segmentation Vision Transformers Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:45.691037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:38.424956Z digest=sha256:7695112f5f37453e500be5aac4fb9a73c61700701704d9b1117c2b1ee1543323

Observation 4b1d9c74-8136-437f-94e8-736a37685071 · outbound

This paper cites Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts.

Native Segmentation Vision Transformers Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:38.530639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:38.530639Z digest=sha256:4cb3e15ae45b05ca7e292c7c5e449ab92028557410df73f7ce71734f79371f0c

Observation 7663898a-73b5-447d-912e-d66af4100685 · outbound

This paper cites Redcaps: Web-crawled image-text data created by the people, for the people.

Native Segmentation Vision Transformers Redcaps: Web-crawled image-text data created by the people, for the people

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:45.509767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:38.622568Z digest=sha256:996ee7e017a06cbced2ad7a93d184700e180ae0cdf6e2e96bcb9e3b10a016933

Observation 421a5750-7cb4-45b2-844d-c2068675156c · outbound

This paper cites Learning to generate text-grounded mask for open-world semantic segmentation from only image-text pairs.

Native Segmentation Vision Transformers Learning to generate text-grounded mask for open-world semantic segmentation from only image-text pairs

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:45.310461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:38.698678Z digest=sha256:6d18881a73e566b8605fbebf6c1cc6d7cefe5c8db65606ba646f29ea52ad1722

Observation e6c0e57e-0842-4104-93aa-39c716c4ff09 · outbound

This paper cites Everingham, L.

Native Segmentation Vision Transformers Everingham, L

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:45.134327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:38.790705Z digest=sha256:8ba16462706272ac1ec967bba516ef53eb46536773b04d50c92874c7c7d32d4a

Observation 72274a1a-d26c-4238-8e16-a879ee05c245 · outbound

This paper cites The role of context for object detection and semantic segmentation in the wild.

Native Segmentation Vision Transformers The role of context for object detection and semantic segmentation in the wild

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:44.877729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:38.870638Z digest=sha256:1e39e08875598a0ec8de5aff3bcf456c23cfbdeda40d514409b7b80b868acb9d

Observation 12ff2374-85bd-42dc-a2a2-3e24c9307cc2 · outbound

This paper cites an unresolved cited work.

Native Segmentation Vision Transformers Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:55:44.748497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:38.933552Z digest=sha256:9a0773d6e1cb890774b614eac9918d0016a582d91a95d3128ab83091362ed6b3

Observation 0863d82c-10c4-4e67-9774-4d4e6d0519a9 · outbound

This paper cites Coco-stuff: Thing and stuff classes in context.

Native Segmentation Vision Transformers Coco-stuff: Thing and stuff classes in context

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:44.550601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:39.012943Z digest=sha256:5488ddfae2c3fda6348f144f0fab7b79b9654432fff5493aa74a9b5906c95f11

Observation 45573f47-e3d0-44e4-acab-46e74982d378 · outbound

This paper cites Cordts, M.

Native Segmentation Vision Transformers Cordts, M

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:39.058264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:39.058264Z digest=sha256:613cf5ba06aec293cfc65ea9ce8e51e391282eeecc3c0259779ce5123b459335

Observation 971d4e1d-0360-4075-8853-572bf21cf678 · outbound

This paper cites Open- world semantic segmentation via contrasting and clustering vision-language embedding.

Native Segmentation Vision Transformers Open- world semantic segmentation via contrasting and clustering vision-language embedding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:44.340556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:39.128618Z digest=sha256:ad403afaeec9f205c286fc171a2ec82a5e29eb73cbe2e05a620ae0673614613c

Observation c5829d93-cf5c-4493-a113-a528c55e49b4 · outbound

This paper cites Segclip: Patch aggregation with learnable centers for open-vocabulary semantic segmentation.

Native Segmentation Vision Transformers Segclip: Patch aggregation with learnable centers for open-vocabulary semantic segmentation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:44.151077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:39.200859Z digest=sha256:c6d3e9adc1024a97950e18585a8c7d9b9dcb3acc754859dec8aee08a1f706b2c

Observation e922457d-570b-402d-829f-1ed192d4d53b · outbound

This paper cites Image-text co-decomposition for text-supervised semantic segmentation.

Native Segmentation Vision Transformers Image-text co-decomposition for text-supervised semantic segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:43.994656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:39.283083Z digest=sha256:cf4d6265d37ae0c5f845b51b6a35dad97538f2e0cd4812d70d64a0d843a4b15c

Observation fab32581-2094-4638-8719-105837d4e6f9 · outbound

This paper cites Viewco: Discovering text-supervised segmentation masks via multi-view semantic consistency.

Native Segmentation Vision Transformers Viewco: Discovering text-supervised segmentation masks via multi-view semantic consistency

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:43.799254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:39.361768Z digest=sha256:366f53efb58373f886a24b8787db4b2176c01afb0483927dba2c8200822b12b2

Observation 8fc88444-f9a1-41e5-9db4-72413e568c66 · outbound

This paper cites Rewrite caption semantics: Bridging semantic gaps for language-supervised semantic segmentation.

Native Segmentation Vision Transformers Rewrite caption semantics: Bridging semantic gaps for language-supervised semantic segmentation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:43.625993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:39.421485Z digest=sha256:d5d11a98a851e941ed3466ea0643d7de3464ad811243f0293de3ac65c113e2e6

Observation fc8b093b-d0cd-47b1-a17d-aabcc22eb2f9 · outbound

This paper cites Uncovering prototypical knowledge for weakly open-vocabulary semantic segmentation.

Native Segmentation Vision Transformers Uncovering prototypical knowledge for weakly open-vocabulary semantic segmentation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:43.458392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:39.494191Z digest=sha256:f54096c8fdba38392f634a364b0646851b97c6f35208fe0769570ffe17aa4688

Observation 6ca380dd-fdbd-4b8c-ad63-df2fcdf1a52a · outbound

This paper cites Efficient inference in fully connected crfs with gaussian edge potentials.

Native Segmentation Vision Transformers Efficient inference in fully connected crfs with gaussian edge potentials

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:43.313744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:39.625717Z digest=sha256:665c66baeba8e7eeda9d5bf07766daa95f1de1f0ec5b267fdae12c267ca44be5

Observation 10d3c55d-06db-4062-bd14-7ef921ee45ea · outbound

This paper cites Single-stage semantic segmentation from image labels.

Native Segmentation Vision Transformers Single-stage semantic segmentation from image labels

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:43.172338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:39.659409Z digest=sha256:4b289ee006095087a1067f7fa79c469442c5685aaf95a6574c0a179173d96454

Observation c961662f-da5d-4207-8db6-5a18fcc46496 · outbound

This paper cites Vmamba: Visual state space model.

Native Segmentation Vision Transformers Vmamba: Visual state space model

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:43.049455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:39.699866Z digest=sha256:3f4d6db258ee213f38b5165b383a4b62148f66742c466d285f5fbcbcc6f638ca

Observation dec6cc42-45af-4d31-ac73-3f8f589fe156 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Native Segmentation Vision Transformers Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:42.871618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:39.776032Z digest=sha256:d2e2d6ab79e0c9b2b91eac3fa80e37116d9bfb57981eb6aa55cef3bb4cb8a18f

Observation 67f73f93-4e98-4907-9f6a-301595fefb4f · outbound

This paper cites Panoptic feature pyramid networks.

Native Segmentation Vision Transformers Panoptic feature pyramid networks

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:42.750426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:39.859502Z digest=sha256:67ee316e6b7c04430e78a3296b573ab7cb39a9509d77c9491507fcad6ce07b36

Observation d62ce440-cb23-407d-a27d-786cc8396158 · outbound

This paper cites Segmenter: Transformer for semantic segmentation.

Native Segmentation Vision Transformers Segmenter: Transformer for semantic segmentation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:42.618951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:39.938874Z digest=sha256:89e8795a7b52bbaed4af8c6c610c67deac04cae2c053f78bee5b30e21ab2768c

Observation 40dc5a72-0625-40f4-8cc7-38b435006774 · outbound

This paper cites an unresolved cited work.

Native Segmentation Vision Transformers Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:55:42.424557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:39.973272Z digest=sha256:cfc99a1b94e51557b23ca51e5209e3e0f9230a863376db4fdd32169d27249da6

Observation a3bbe0f1-d2f1-4079-87cf-a4b174ad51be · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

Native Segmentation Vision Transformers Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:42.334089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:40.084480Z digest=sha256:13d1003171d7c1cdd9219f7e05b26fb1faeb44dde32f79ca1661518ce0fab9b3

Observation 1f78a695-9c15-4dbe-9b6a-30f1560e0474 · outbound

This paper cites Self-attention with relative position rep- resentations.

Native Segmentation Vision Transformers Self-attention with relative position rep- resentations

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:42.192849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:40.167748Z digest=sha256:1af01ab007f0d8589c14c9bfa03f5ea81476abfe6941845134e209ef8d21d571

Observation 8c28d6ac-aa0c-4cb7-813d-7735c818d99a · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Native Segmentation Vision Transformers Roformer: Enhanced transformer with rotary position embedding

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:42.054391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:40.218252Z digest=sha256:faa8649ae5aa13ae9f0f4104d8ace814f0e3a10f40cec8c0ab9e954d4b32d90d

Observation 437ba9f9-51cf-4272-8968-749e3242de6f · outbound

This paper cites Rotary position embedding for vision transformer.

Native Segmentation Vision Transformers Rotary position embedding for vision transformer

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:41.934310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:40.261863Z digest=sha256:48cea0fa951286bb4112840b9c1cfcaca51d1d9944096c0144f5ebefa07e7843

Observation bdfafc2b-bed4-4ea6-a779-5aaf230de942 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

Native Segmentation Vision Transformers Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:41.770603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:40.305574Z digest=sha256:4229ca8a30a2e22beb335d8b0a957fc5ea22946d95754c697b0aa87abe37e250

Observation 3f1e26f5-c66c-415e-bb6b-d6518d192254 · outbound

This paper cites Faster neighborhood attention: Reducing the o(n^2) cost of self attention at the threadblock level.

Native Segmentation Vision Transformers Faster neighborhood attention: Reducing the o(n^2) cost of self attention at the threadblock level

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:41.619633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:40.399294Z digest=sha256:85f1b144eb4aa0ed054829e29d0bfde31b2007637cdd8f536fb9aadfb06eff97

Observation 14e5460f-1f9e-4ca8-8d8e-684f2faea1de · outbound

This paper cites Training data-efficient image transformers; distillation through attention.

Native Segmentation Vision Transformers Training data-efficient image transformers; distillation through attention

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:41.448153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:40.479785Z digest=sha256:1fd8e9dcc9b9c7563f51b87f6b9c73303dc6d2ba37a39e1e7c94658869eaa8a6

Observation 8a207c97-2d2a-41fc-b4af-80f60be67ba7 · outbound

This paper cites Weinberger.

Native Segmentation Vision Transformers Weinberger

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:41.352789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:40.578280Z digest=sha256:e05ae0c230ba222fc27e4ffbffc28fbd404ca7b33143b1953fbac46da43aec73

Observation bfa2a9c9-10fd-4b9b-a662-2997cd397c59 · outbound

This paper cites Empirically, we observe a negligible decrease of downstream classification and segmentation performance, and a significant increase in speed.

Native Segmentation Vision Transformers Empirically, we observe a negligible decrease of downstream classification and segmentation performance, and a significant increase in speed

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:41.181944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:40.651143Z digest=sha256:95d0643af5deed9e5a3db5d6b798566f9853bd8b7dac16a1c5a7b8636fe95a8c

Observation 92b5b608-2b35-4808-89b9-bab8545c1252 · outbound

This paper cites An image of {CLASS }.

Native Segmentation Vision Transformers An image of {CLASS }

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:41.063393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:40.725910Z digest=sha256:01622794fe9598c60e1a0e30786170fa2e80c0fb6362d18eb0524c8489a0fde0

Observation e32f542a-07bc-4bab-b9e8-6397532cc881 · outbound

This paper cites an unresolved cited work.

Native Segmentation Vision Transformers Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:55:40.952338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:55:40.776768Z digest=sha256:4fe1755598df08dbd669b45e7b7028274e5ce5fc93ae3bd89dace7a935675d95

Pith citing papers

Observation ade4aadf-4816-44a2-9198-566b2ec09f15 · inbound

Elastic Attention Cores for Scalable Vision Transformers cites this paper.

Elastic Attention Cores for Scalable Vision Transformers Native Segmentation Vision Transformers

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:22.541489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T06:02:40.158866Z digest=sha256:e040df38698881283a566d141490df09a0221b389dce5db66e12e24373ee1417