Pith. sign in

Paper Citation Record · LEDGER

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation

As of 19 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2506.11131.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.11131 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:01:23.882872Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1e509b39-0717-4dce-a7b4-5ced0adbec71 · outbound

This paper cites Zerowaste dataset: To- wards deformable object segmentation in cluttered scenes.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Zerowaste dataset: To- wards deformable object segmentation in cluttered scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.746199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.694098Z digest=sha256:d3e6b81748bf76db29fd8a691b60507f5d0de3aa626134f7ad94550549b15fd6

Observation 28594f12-b209-49ab-a85a-cfe51ce3f4f5 · outbound

This paper cites Token Merging: Your ViT But Faster.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Token Merging: Your ViT But Faster

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.698370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.698370Z digest=sha256:2887b464d3b2d8a76c8a24cc5c7b384eff2b3aa38ad4b2ae91a61fdde4d7201b

Observation c86c28c2-8ee7-40b4-9a06-65874c14c5fd · outbound

This paper cites The eccentricity effect: Target eccentricity af- fects performance on conjunction searches.Perception & psychophysics, 57:1241–1261, 1995.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation The eccentricity effect: Target eccentricity af- fects performance on conjunction searches.Perception & psychophysics, 57:1241–1261, 1995

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.694667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.702224Z digest=sha256:92fb2afe0edac4e488e072c3e1ea82c643514c787d2d00389e27b522189d5f82

Observation e2183c98-8fcb-4a6c-bb8d-b4cd06db1fb2 · outbound

This paper cites Pelk: Parameter-efficient large kernel con- vnets with peripheral convolution.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Pelk: Parameter-efficient large kernel con- vnets with peripheral convolution

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.642470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.705985Z digest=sha256:e9ee622e3d5d456f2aebef2e00e0dd48e084ea8cb10c2d47350fb2eb135b6c29

Observation 55063883-09bd-41a8-89dd-fafbd792858a · outbound

This paper cites Diffrate: Differentiable compression rate for efficient vision transformers.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Diffrate: Differentiable compression rate for efficient vision transformers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.627947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.709324Z digest=sha256:989ecc02d81d08079ec4149fab6b5f56a89cbd06ec1103e0f3b5f9dbdb992e0f

Observation 9b378b04-b665-45ab-af7e-f1ae93f46f5b · outbound

This paper cites The cityscapes dataset for semantic urban scene understanding.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation The cityscapes dataset for semantic urban scene understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.712548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.712548Z digest=sha256:78a8c131eddf564443ba8834d6d48d05bfbbc075350bab1e6dbf93d0d28ddbaa

Observation 8327ace9-0161-4fb0-914d-af88c483025c · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and chal- lenges for epic-kitchens-100.International Journal of Com- puter Vision, pages 1–23, 2022.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Rescaling egocentric vision: Collection, pipeline and chal- lenges for epic-kitchens-100.International Journal of Com- puter Vision, pages 1–23, 2022

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.607128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.716082Z digest=sha256:9f33137ba2c6b1271161a11ea8dd58af71b19049f780e7a7d0117e99be5c76d3

Observation 55363477-3f56-4009-bd58-67cd24293673 · outbound

This paper cites Vision Transformers Need Registers.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Vision Transformers Need Registers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.719371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.719371Z digest=sha256:145c9165b233569a4ff0485c10c1658b56ea65fb4e45159d0e22f654c7d51642

Observation eaf99e91-6ff2-42c3-a664-fb632d5f4745 · outbound

This paper cites Epic-kitchens visor benchmark: Video segmenta- tions and object relations.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Epic-kitchens visor benchmark: Video segmenta- tions and object relations

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.593588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.722720Z digest=sha256:5a623bb2ff57701833b3ccebe24023931ec584c220410c1865a30bff229818d3

Observation 0f531688-f6f3-4537-9755-c3accd759b1f · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Imagenet: A large-scale hierarchical image database

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.580040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.725866Z digest=sha256:123b13c1b7cc23b038cef210221e1bccd7b72abb99c34600f1e9ef51592d08c5

Observation 5bd0c6c2-48f7-406f-8e90-cb51c855634a · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.728704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.728704Z digest=sha256:1569ca0cb0432f597481c9e71a4e8078796058bf619be6a1cb0dd1eced3d0faf

Observation 72ea821b-b086-4d6d-b65a-2f06aad9f9ab · outbound

This paper cites Project Aria: A New Tool for Egocentric Multi-Modal AI Research.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Project Aria: A New Tool for Egocentric Multi-Modal AI Research

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.732463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.732463Z digest=sha256:81bfb8c65f4e6f1a139d2c52c99b7dc46aa43ddf8b7a7d0aa01b68308b91339a

Observation 2b6c4ab8-f20c-40e8-8e11-08ac62ec1bf0 · outbound

This paper cites Adaptive token sampling for efficient vision transformers.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Adaptive token sampling for efficient vision transformers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.554245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.735627Z digest=sha256:886447fddfd7d38ae6fe6cbd8ccb5aa829e7a30594351fad225fb3482c1c0c81

Observation 21c265b8-1918-4a76-8f95-917f683ff375 · outbound

This paper cites Instance segmen- tation for autonomous log grasping in forestry operations.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Instance segmen- tation for autonomous log grasping in forestry operations

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.515908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.738525Z digest=sha256:dcd5dd4944ca24c7fa8468f987bf4bc817403d087e0388188f9cbe7735cad7fb

Observation 99a98a1c-017e-41cd-a25a-d4f760767f8c · outbound

This paper cites Deep residual learning for image recognition.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Deep residual learning for image recognition

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.463612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.741588Z digest=sha256:89623e323bef29538e2de354136dc4c0e9985fdac85a0da7932ec6f950646e35

Observation 18b36cfc-202c-4303-b830-5c8676a3cd26 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Masked autoencoders are scalable vision learners

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.445343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.744538Z digest=sha256:183a7aeed3a29930e7e7dc13ce510c4c2710df613cf12d87bea300e80f2fcccb

Observation 73d796ab-7a26-4cb7-a8bc-8af8e40e5868 · outbound

This paper cites Bytes Are All You Need: Transformers Operating Directly On File Bytes.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Bytes Are All You Need: Transformers Operating Directly On File Bytes

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.747349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.747349Z digest=sha256:86c8a306628ebc086759b135d46663c6939ca1ffe840f06e526ca8e21c401cdb

Observation 40a1255e-c691-4dd0-81b2-330d8e2c0bec · outbound

This paper cites FoveaTer: Foveated Transformer for Image Classification.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation FoveaTer: Foveated Transformer for Image Classification

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.751062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.751062Z digest=sha256:fa73dc7daf65efc9a1a65a9c2391b738c6cd07ea44c847b66e56532505773072

Observation 0fa90366-b06b-4c18-999a-852cc05fc66c · outbound

This paper cites Scaling Laws for Neural Language Models.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Scaling Laws for Neural Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.754450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.754450Z digest=sha256:84c6fd3f5f640af5bc0a0cef59f66899df478683f5b0b1f826dc4ce752f19a62

Observation 8e62f6ee-dc64-4cb4-96a4-a8b2d340878b · outbound

This paper cites Segment any- thing.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Segment any- thing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.433749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.758293Z digest=sha256:cd8ce65861be194103411e71dff7731742cd10b1cffb4b55e9da465ab4b05177

Observation c7db607a-ee3c-4385-8c74-e1c25ba8a50b · outbound

This paper cites Spvit: Enabling faster vision transformers via latency-aware soft token pruning.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Spvit: Enabling faster vision transformers via latency-aware soft token pruning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.422770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.761078Z digest=sha256:0b507b999d7121f99bdb90327d40cd731097c4b9521fcdbeb4c24bdf92e6ff5c

Observation 9ca957d8-7dc4-43bd-9854-dd88108d4785 · outbound

This paper cites GazeGPT: Augmenting Human Capabilities using Gaze-contingent Contextual AI for Smart Eyewear.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation GazeGPT: Augmenting Human Capabilities using Gaze-contingent Contextual AI for Smart Eyewear

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.764359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.764359Z digest=sha256:8e010975fe7ec39c609ffabcfa64dfe4690d217738268afefbf9bc3376061359

Observation 590a6293-dc25-4f1d-9958-f665e0569fb0 · outbound

This paper cites Imagenet classification with deep convolutional neural net- works.NeurIPS, 25, 2012.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Imagenet classification with deep convolutional neural net- works.NeurIPS, 25, 2012

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.411540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.767641Z digest=sha256:0c18aae9d18a4e56b1c1a88adf4ed436f08989695f848cd1eb53faee24e05c33

Observation 63c6b7bf-c21b-49ac-a983-781244081ff1 · outbound

This paper cites Microsoft coco: Common objects in context.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Microsoft coco: Common objects in context

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.771069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.771069Z digest=sha256:fdaaf649721c2e07a9ca288ed83c520abbd6df9645a02069a9cbae4aa09102aa

Observation 36157df9-59b5-4627-a570-b74ea1f749c3 · outbound

This paper cites Efficientvit: Memory efficient vision transformer with cascaded group attention.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Efficientvit: Memory efficient vision transformer with cascaded group attention

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.390481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.774471Z digest=sha256:9333d215b9bab05495cb7450d42e5b56d2b06227df9e46359acc60b286804917

Observation 084b4e53-0296-46e9-af71-b67bc7baa978 · outbound

This paper cites Nymeria: A Massive Collection of Multimodal Egocentric Daily Motion in the Wild.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Nymeria: A Massive Collection of Multimodal Egocentric Daily Motion in the Wild

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.777693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.777693Z digest=sha256:d19deee5929141f5a52fb2af46af8bb78ab3f30b5b6672d1ff4169586a4116df

Observation 788a1717-40dd-4ee4-8aa0-f37e6cdbb4a8 · outbound

This paper cites Token Pooling in Vision Transformers.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Token Pooling in Vision Transformers

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.781218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.781218Z digest=sha256:9c0942b1f792aa4fef78085e976cdea834272e9e3bc9847757ce715ccc4d7d4a

Observation a7e9de63-7a8c-483d-a0e1-f877df98dc81 · outbound

This paper cites Adavit: Adaptive vision transformers for efficient image recognition.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Adavit: Adaptive vision transformers for efficient image recognition

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.377985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.785126Z digest=sha256:6ef09f99a91d0c096f35cebd451cf9090260c76d54e31a1ea5d393e632cab5bd

Observation 4c17ce08-a961-4e2a-b3ca-24f2b5bde1e1 · outbound

This paper cites Peripheral vision transformer.NeurIPS, 35:32097–32111,.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Peripheral vision transformer.NeurIPS, 35:32097–32111,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.365746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.788274Z digest=sha256:26c8ec16907a8ce2a185c761b8ebac7225a89976498b84001cfc2e0d5326d242

Observation f56ad5c0-676c-4893-aac9-98ab94bf32b6 · outbound

This paper cites Finely-grained annotated datasets for image-based plant phenotyping.Pattern recognition letters, 81:80–89, 2016.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Finely-grained annotated datasets for image-based plant phenotyping.Pattern recognition letters, 81:80–89, 2016

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.354346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.791261Z digest=sha256:0d7105632545c8e0adeec59a1f90723144719eab3fba6e612cf918257e5441e0

Observation 1235dc9c-c215-4930-968f-9f7b0e3a0247 · outbound

This paper cites Rgb no more: Minimally- decoded jpeg vision transformers.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Rgb no more: Minimally- decoded jpeg vision transformers

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.341396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.794455Z digest=sha256:d293ef2a46e70f65e00ba2b79cbb9e22301874dc9291504927736858bc0c906b

Observation cbf95ed2-4d45-4c90-a0bd-b8014a99c4cd · outbound

This paper cites Dynamicvit: Efficient vision transformers with dynamic token sparsification.NeurIPS, 34:13937–13949, 2021.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Dynamicvit: Efficient vision transformers with dynamic token sparsification.NeurIPS, 34:13937–13949, 2021

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.329601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.797553Z digest=sha256:6badddb31e2093521a1ce8b564728700e6ad067814e9123927f03915ecf5290d

Observation 69d40b15-99d7-4f63-8ae8-d2561eba0191 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation SAM 2: Segment Anything in Images and Videos

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.800874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.800874Z digest=sha256:2c29913029b80b694b9c0a274e196801cf53d382b9c754575213a29775d4c1df

Observation c65e594d-c154-445b-a67a-def22ddd5241 · outbound

This paper cites Learning to Merge Tokens in Vision Transformers.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Learning to Merge Tokens in Vision Transformers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.804525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.804525Z digest=sha256:291a8b2d06ce087ce8b627d6da9e3a5f9c4f7cce7031e84cb4897a0291751836

Observation 7ccae9cd-45a8-41a1-a6a8-b6cb26b7a985 · outbound

This paper cites CP-ViT: Cascade Vision Transformer Pruning via Progressive Sparsity Prediction.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation CP-ViT: Cascade Vision Transformer Pruning via Progressive Sparsity Prediction

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.807850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.807850Z digest=sha256:e7494effc74bd285031f0c27b0935ef08e71326509fa7d72136ba4953fdd2e3b

Observation e8f10b38-f447-44eb-968a-3d3eaf864169 · outbound

This paper cites On Efficient Variants of Segment Anything Model: A Survey.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation On Efficient Variants of Segment Anything Model: A Survey

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.811323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.811323Z digest=sha256:de97c5e6241ce3b354e6ec7aa1eff6397c1a54c29604edbe1567606b71a64128

Observation d23f2a49-6e4d-41b8-84f9-aba653abed04 · outbound

This paper cites NDD20: A large-scale few-shot dolphin dataset for coarse and fine-grained categorisation.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation NDD20: A large-scale few-shot dolphin dataset for coarse and fine-grained categorisation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.814417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.814417Z digest=sha256:a39fd630325faaa01d685974f692854b4f421aad07d0cd333ae49cb23a6864f2

Observation 43bd1eec-cba5-4ced-a584-1a7f911a7b8d · outbound

This paper cites Neural discrete representation learning.NeurIPS, 30, 2017.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Neural discrete representation learning.NeurIPS, 30, 2017

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.818098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.818098Z digest=sha256:2c8d57c9516b4f57236f1cbb6379532481238bad81d069ff36a49a28dcde29d8

Observation b4ca801e-802b-431c-a17c-b8c84bc54fda · outbound

This paper cites SqueezeSAM: User friendly mobile interactive segmentation.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation SqueezeSAM: User friendly mobile interactive segmentation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.821808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.821808Z digest=sha256:669c6a934e4c03decd9f51c8913cbb97ad58c12255b8c9f2de427541bfbdc21f

Observation 39280563-a8b6-411a-928b-0b8736568912 · outbound

This paper cites Efficientsam: Leveraged masked image pretraining for efficient segment anything.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Efficientsam: Leveraged masked image pretraining for efficient segment anything

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.309971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.825145Z digest=sha256:b24f458fa8957a06f565bfdc7e6520259f4cffa767e84ef9065f55b3c3bd6d12

Observation 4fb11bf9-fcfe-488a-bb6c-46e94179d6f8 · outbound

This paper cites ElasticTok: Adaptive Tokenization for Image and Video.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation ElasticTok: Adaptive Tokenization for Image and Video

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.828275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.828275Z digest=sha256:ccf51a9c0d9c7a2a9b1db556704fea265f1ebbe9ceec3ab96a8f441845d5aa05

Observation d5c30b82-a863-4e20-8dfa-befadddf50fa · outbound

This paper cites A-vit: Adaptive tokens for efficient vision transformer.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation A-vit: Adaptive tokens for efficient vision transformer

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.299277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.831259Z digest=sha256:fe9d080e42352f40319c323812c68d08b0649516eee001f6e8e9d02a9f254498

Observation 0e845d06-82ed-4ee2-8822-c956a82a55ec · outbound

This paper cites Woodscape: A multi-task, multi-camera fisheye dataset for autonomous driving.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Woodscape: A multi-task, multi-camera fisheye dataset for autonomous driving

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.288980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.834300Z digest=sha256:7619a82aa2a1b11d3d1d017a9991ee74fafe633903e8ade04b513d5002401f11

Observation 51258c3e-8c81-49a3-8ae8-791671ac755e · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.837602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.837602Z digest=sha256:c14c554f2fc39c53cf9f570d83f3c4cee506fcd6027868239919336d4ccbe023

Observation 1406d1ef-0d27-4c18-a42b-e0873baf7cbf · outbound

This paper cites An Image is Worth 32 Tokens for Reconstruction and Generation.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation An Image is Worth 32 Tokens for Reconstruction and Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.840987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.840987Z digest=sha256:4673f2fed7cf3de2a2509d4fcf87e126f3517f362e092e7f5b781ff0c8b31e05

Observation c971f6c8-29c5-4407-8401-1fe63677d38d · outbound

This paper cites Faster Segment Anything: Towards Lightweight SAM for Mobile Applications.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Faster Segment Anything: Towards Lightweight SAM for Mobile Applications

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.844273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.844273Z digest=sha256:81b3386df94352453b312d613e35cb903d35706d4220d16269176156ff1af3ea

Observation 09b06f40-1aee-42b8-9c5a-6f73400cc95c · outbound

This paper cites Fine-grained egocentric hand-object segmentation: Dataset, model, and applications.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Fine-grained egocentric hand-object segmentation: Dataset, model, and applications

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.278276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.847206Z digest=sha256:8b41261a2c20aaac87c099064f086241674136568f8c7802e4e8ea1a91cdde7d

Observation 13e08c05-b92e-4eb8-a142-e51cfd207e2d · outbound

This paper cites Efficientvit-sam: Accelerated segment anything model without performance loss.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Efficientvit-sam: Accelerated segment anything model without performance loss

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.267838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.850081Z digest=sha256:347ef0e6040a3dd0cacf859d5986b7d5c95e56ced3a46b7eb98ae6802062772d

Observation e57b58ee-d1ee-4a3b-bb66-c6f4e81d541b · outbound

This paper cites Fast Segment Anything.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Fast Segment Anything

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.853152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.853152Z digest=sha256:ec47f0682185c5deb678ca051c6155d1441637c463a75558a373abcba828d719

Observation 7ce1e0d3-1524-42d2-bd27-260cfd55224e · outbound

This paper cites Semantic under- standing of scenes through the ade20k dataset.IJCV, 127: 302–321, 2019.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Semantic under- standing of scenes through the ade20k dataset.IJCV, 127: 302–321, 2019

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.256436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.856235Z digest=sha256:4699d41b7c222d0a68e1a113dd0b2b684cc180594b708a3e0b80c4dda8c76db8

Observation 4f101be8-ea3f-41fc-980f-1c44a8ac2330 · outbound

This paper cites EdgeSAM: Prompt-In-the-Loop Distillation for SAM.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation EdgeSAM: Prompt-In-the-Loop Distillation for SAM

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:23.859111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:23.859111Z digest=sha256:265434af871e5fdda2dfe2e412cf06643c47f45413d55438ba88a60614127d22

Observation 9a6763cf-b739-4160-aff5-205b06f3fbe4 · outbound

This paper cites MAE Pre-training We pre-trained our foveated image encoders using MAE pre-training for 500K iterations.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation MAE Pre-training We pre-trained our foveated image encoders using MAE pre-training for 500K iterations

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.246207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.862624Z digest=sha256:c991f326893f4f6a2e241ee86e3de9302968795a8b75b608b26702e2cb97390f

Observation ed16bc5d-091d-4f38-bd2f-fd6573629003 · outbound

This paper cites Here we give a more formal definition of the parameterization of such a pattern.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Here we give a more formal definition of the parameterization of such a pattern

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.235663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.866031Z digest=sha256:e6c6a89061b6f236c3acf165e895174d0dc9cf9878ea0ab156fa9eaca5e436e6

Observation 47069a74-dd08-4d4f-8a35-60ad7cff6bd0 · outbound

This paper cites We plot the training loss curves in Figure 12.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation We plot the training loss curves in Figure 12

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.223472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.870145Z digest=sha256:7f6bb52f66b43427413c5111d420e820a96e26f6e316b35e482032d3db73bf43

Observation e45f3356-3b77-40e8-bd38-290c2a6498fd · outbound

This paper cites Our foveation patterns exists in a high-dimensional design space, and each new pattern requires its own MAE pre-training.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Our foveation patterns exists in a high-dimensional design space, and each new pattern requires its own MAE pre-training

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.211826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.873570Z digest=sha256:04e330fcfdc22efa267e5d4800d1fc3049a0aebc26727f0825cb2fed886e785a

Observation 813448ba-c64e-4019-a5c7-c23e2205b705 · outbound

This paper cites an unresolved cited work.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:24.200458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.876636Z digest=sha256:81dd5ab804e3f6347af609d30ee80de30f9c686810d3385d07da3b5263f9bb51

Observation 399a220f-faf9-4e66-be1f-94a640639ef8 · outbound

This paper cites an unresolved cited work.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:24.190238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.879791Z digest=sha256:6ec018dfb8de23dcef22f5e29dfc4bb4028bfa5bbb96b03a84d15957616d940b

Observation 51feced0-379a-438d-ba46-36a366146a89 · outbound

This paper cites in computing FLOP counts for transformer architectures (c.f.

Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation in computing FLOP counts for transformer architectures (c.f

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:24.179128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:23.882872Z digest=sha256:0adc14637d66faa4ce75c2b84f8a521e63ea9ad9da387af6985d98219003063d

Pith citing papers

No inbound Pith citation observations are available.