Pith. sign in

Paper Citation Record · LEDGER

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation

As of 11 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2502.02763.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02763 v3

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:19:23.191967Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T00:18:06.989944Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T00:19:46.820702Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aef2c08a-e2ef-4762-bcaf-da5d24a04762 · outbound

This paper cites Focal sparse convolutional networks for 3d object detection.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Focal sparse convolutional networks for 3d object detection

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.532068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.079589Z digest=sha256:544ff62fcb3410a3ec207bc0129bee8295c6487acd7a16ebb7f6751624f36211

Observation 6b9a0d16-b217-4c80-a1e9-e18b7178aeaf · outbound

This paper cites Deformable convolutional networks.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Deformable convolutional networks

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.516609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.084083Z digest=sha256:101ecf0338e59e476197e99a1a21fddf3ffd9e6db53bc26fc8a12de3ca4050a2

Observation e692c0da-2b99-498e-ab61-ddf0e784fd93 · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Objaverse: A universe of annotated 3d objects

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.505803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.088072Z digest=sha256:32e1817132777a0dec7bea2c9fc97fc1802dc0e2c01f7500c53beab1a33e8b56

Observation 5a218b3b-ef72-4de6-a4a2-0ea29ab5a7e2 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T11:19:23.091877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:19:23.091877Z digest=sha256:d2050591bd54d2dc45767f9ec4b211ccf96e3b66f0b9b8bd6d82ff4d7bc4deec

Observation ed9c03ac-82bf-4c07-8931-eda19d4fd0ee · outbound

This paper cites C., and Kipf, T.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation C., and Kipf, T

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.494268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.096191Z digest=sha256:5abc1cb2ea5c2a7215a40ff0b7b8570e4d1011d2dc36b4dfe381e3f8b3c032f8

Observation e355b5a9-61dc-4867-bed8-e3bc0e20e8dd · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Lvis: A dataset for large vocabulary instance segmentation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T11:19:23.099713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:19:23.099713Z digest=sha256:b4f40d37328ad4be8a89c2149cd46b4d1436be2c024114410cae339797c22dcf

Observation 63c4ade8-74ca-41fa-acb4-3271184e0ee0 · outbound

This paper cites Perceiver IO: A General Architecture for Structured Inputs & Outputs.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Perceiver IO: A General Architecture for Structured Inputs & Outputs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T11:19:23.103509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:19:23.103509Z digest=sha256:f5205707b5f5dcbc86f76cc4a4b4042fb350dd8f8394156e9eef02c4c1b1053a

Observation d6085e78-f0da-42b2-80e4-1bc5be682996 · outbound

This paper cites S., Sochenov, A., Leimk \"u hler, T., Okunev, M., Goodall, T., and Rufo, G.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation S., Sochenov, A., Leimk \"u hler, T., Okunev, M., Goodall, T., and Rufo, G

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.477171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.106998Z digest=sha256:4ec46a79fa06d846f902f1791efe2b945f2028ea404ee20baa980a1574c51ac7

Observation 348ee7a9-8c1f-46a1-baf1-ac6afdd9af4c · outbound

This paper cites C., Lo, W.-Y., et al.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation C., Lo, W.-Y., et al

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T11:19:23.110318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:19:23.110318Z digest=sha256:c56972a96c062114ec0fa4f1c13505ecff9ec2d0a3501e370a5c392d72f05480

Observation 2727d70f-9aa7-45b2-9a47-13d1e7b60a63 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T11:19:23.113657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:19:23.113657Z digest=sha256:dbeb683e618e829a0de353dd9ef991b357e39f81d94558c14e7b3d128ba8cf11

Observation 61d094aa-1c0b-476c-a91f-1cb4287a185e · outbound

This paper cites Kitti-360: A novel dataset and benchmarks for urban scene understanding in 2d and 3d.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Kitti-360: A novel dataset and benchmarks for urban scene understanding in 2d and 3d

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.453297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.117269Z digest=sha256:9b10d2370144e9533d379af24f225e0b62d77be5e6a2befb6bea8fae9a7753a3

Observation dfcf102c-27e3-47a9-b365-35c904a1c5fe · outbound

This paper cites an unresolved cited work.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T11:19:23.120823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:19:23.120823Z digest=sha256:b650ade1e292e495e172960a916e3e7ebcb1a906cdc870c79d09cf1bc2772f87

Observation ecbb7336-e589-441a-8320-b7a771ff23a5 · outbound

This paper cites A convnet for the 2020s.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation A convnet for the 2020s

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T11:19:23.124894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:19:23.124894Z digest=sha256:e3b2c028ab2adbebaa477e162cb24208d54c51fd8a91fea18791dfa62a7ab830

Observation 4761056f-94e7-4bdf-86a2-2f0bac12a388 · outbound

This paper cites Object-centric learning with slot attention.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Object-centric learning with slot attention

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T11:19:23.128418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:19:23.128418Z digest=sha256:5575e0536ea7ae2215973d3091cf28e77593e21e607b1ceaa2fd816621f4b140

Observation fa0d062a-7b94-4f7a-95b8-f599a6995fef · outbound

This paper cites Biologically inspired deep learning model for efficient foveal-peripheral vision.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Biologically inspired deep learning model for efficient foveal-peripheral vision

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.423820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.132289Z digest=sha256:094d9fe278966b6ddabb083101ea4f6b5312d65c617b737faeefbe9c703d03a0

Observation 39862b7b-c7af-4362-bf18-b8b4f76699e2 · outbound

This paper cites A., Paczan, N., Webb, R., and Susskind, J.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation A., Paczan, N., Webb, R., and Susskind, J

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.412652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.136215Z digest=sha256:9720744399d2c7c0bd18f7a19e2a58d3f1a17458fd8050e06f87770aa3aa066b

Observation 3de9cd10-298a-4932-87b4-247187451080 · outbound

This paper cites Simple unsupervised object-centric learning for complex and naturalistic videos.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Simple unsupervised object-centric learning for complex and naturalistic videos

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.401449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.140120Z digest=sha256:a9482da7a5cf81c0fc024d22da8042fb806ee530a85ecb15fd44ef2ffc7bb69a

Observation f73b75ad-9853-491b-b4fb-24f6c98c6f8b · outbound

This paper cites Fovea: Foveated image magnification for autonomous navigation.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Fovea: Foveated image magnification for autonomous navigation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.389788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.143775Z digest=sha256:77a1b38a43cdccde046f9f1f5a80adcf6eb50b5cb7641277175ede03fc7879d3

Observation 5ebfc136-c11f-4d3a-8233-7baeeafc4e15 · outbound

This paper cites an unresolved cited work.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:19:23.378406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.147186Z digest=sha256:316b60d80a0d79b276b6478dbb4b78488d9774697bf7e1a8c3c40316113794ae

Observation ad15cafb-baf6-44eb-a9fa-b5ec089c81c3 · outbound

This paper cites an unresolved cited work.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:19:23.366867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.150628Z digest=sha256:74201a7919a095214c261e56b0bf55104bad08aee0cabee984f5a101daebe224

Observation 11f94f2c-7b34-40ac-829c-2a84283a96fc · outbound

This paper cites an unresolved cited work.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:19:23.355602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.154164Z digest=sha256:aacb2618b9f494be2e3b2cccee9fe264fc6183536d6fbabccf0f62992e3be15e

Observation de4422ff-b813-4f9b-8bd6-1d03b448c22e · outbound

This paper cites Tinyvit: Fast pretraining distillation for small vision transformers.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Tinyvit: Fast pretraining distillation for small vision transformers

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.344416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.157736Z digest=sha256:440fe99f38aeea9084f79d798a1685035a228d7285e638147f5ea067ed014bda

Observation b1807004-4610-4e87-9e7d-d8620f8f41f1 · outbound

This paper cites EfficientSAM: Leveraged Masked Image Pretraining for Efficient Segment Anything.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation EfficientSAM: Leveraged Masked Image Pretraining for Efficient Segment Anything

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T11:19:23.161236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:19:23.161236Z digest=sha256:ad8ef0075e59b2722fc95d5a16c3ac41c666fa1e5e3f28dd9476feffe79b67d1

Observation 5022d81f-a49d-4d69-bf07-c58193a110cc · outbound

This paper cites Efficient deformable convnets: Rethinking dynamic and sparse operator for vision applications.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Efficient deformable convnets: Rethinking dynamic and sparse operator for vision applications

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.333517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.165209Z digest=sha256:189b8703493710facdeb2d6b9dff3e65381a47eb3cb501f5b44060a586d25198

Observation 1283680c-c063-4688-ac1a-a4cec288a2b4 · outbound

This paper cites Lape: Layer-adaptive position embedding for vision transformers with independent layer normalization.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Lape: Layer-adaptive position embedding for vision transformers with independent layer normalization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.321516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.168897Z digest=sha256:72d24d096e59a1a198d4fd642727ffc8ebfe05f02af2dd5f923817141b66ea98

Observation 572d0087-8a8b-49e8-9367-5a618df5a46b · outbound

This paper cites Hdri haven, 2016.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Hdri haven, 2016

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.309605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.172336Z digest=sha256:4a70c2256fcfb13d937f15c437c449d84e7e6011ff361f531eb0916bd07261b6

Observation 3526f2cc-6731-4960-a531-8761ca17a5f1 · outbound

This paper cites Object-centric learning for real-world videos by predicting temporal feature similarities.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Object-centric learning for real-world videos by predicting temporal feature similarities

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.297696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.176069Z digest=sha256:108409c4e618b1e802a1150fd329a3a6f954a44d21068a21fc3b7dd921b6074f

Observation 456ce933-29a1-431c-87ba-1339b93fbc60 · outbound

This paper cites Faster Segment Anything: Towards Lightweight SAM for Mobile Applications.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Faster Segment Anything: Towards Lightweight SAM for Mobile Applications

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T11:19:23.179663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:19:23.179663Z digest=sha256:7cef2a479f49e3a52be4fc10f6cc7cb37c647c3e1d653c2e1e9565869f8574f3

Observation ecee854f-05e6-4dcc-8e8e-7774c690e1d5 · outbound

This paper cites Fast segment anything, 2023.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Fast segment anything, 2023

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.286527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.183512Z digest=sha256:926cc6e72902d848a50c3fedcb8a939149ba239204cde240831e096aa9957734

Observation 698f2ba0-d7ed-4fcd-9f6d-e244b795f22b · outbound

This paper cites Deformable convnets v2: More deformable, better results.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation Deformable convnets v2: More deformable, better results

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:19:23.275476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-09T11:19:23.188345Z digest=sha256:d9018661d49fcc9b0f8fc29c307d1e2e979a0ba209d5ea7835b2636faba7a373

Observation a8a9c764-eca2-407d-87bc-0a74a9bf7549 · outbound

This paper cites write newline.

Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation write newline

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T11:19:23.191967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:19:23.191967Z digest=sha256:479673dd938bf9759d24eb77940695bb59a7d7ab632f1b07b3ad89d338c8da9d

Pith citing papers

Observation be67138d-9cdd-4c25-973e-ac8abc8f5a1d · inbound

Self-supervised pretraining for an iterative image size agnostic vision transformer cites this paper.

Self-supervised pretraining for an iterative image size agnostic vision transformer Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-13T01:17:40.500289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T00:18:06.989944Z digest=sha256:ad0af7cca173a3c5b03ea10f204f7416dd59bf91ec958f32717779464990d327