Pith. sign in

Paper Citation Record · LEDGER

Object-Centric Vision Token Pruning for Vision Language Models

As of 4 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2511.20439.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.20439 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:19:57.764281Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7bdb75b9-0d88-489e-adf9-e36e7ed32e77 · outbound

This paper cites Qwen2.5-VL Technical Report.

Object-Centric Vision Token Pruning for Vision Language Models Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.638438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.638438Z digest=sha256:b3504fb06a135f057806633c224e02d79b03ba32804713ba856cc8df397a9e11

Observation 4f0d3d5d-0777-4283-9fe4-89b8b9592b11 · outbound

This paper cites Invariant Slot Attention: Object Discovery with Slot- Centric Reference Frames.

Object-Centric Vision Token Pruning for Vision Language Models Invariant Slot Attention: Object Discovery with Slot- Centric Reference Frames

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.643279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.643279Z digest=sha256:6946dcb8d380750541ad2a590e2229ef5cd7fc8c7352e828f789878e4747dfe3

Observation 76d49dd9-2e87-49fb-a2ea-e98103c4d516 · outbound

This paper cites Token merging: Your vit but faster.

Object-Centric Vision Token Pruning for Vision Language Models Token merging: Your vit but faster

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.647003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.647003Z digest=sha256:17d0dd431b244289439487d851f1129de2cb33178234a7258110435a5b9db2eb

Observation a3efce13-acb8-4973-aafb-ba9f97e94bbf · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

Object-Centric Vision Token Pruning for Vision Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.650849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.650849Z digest=sha256:81ebb9813e2a6baa01474e0c92258fba50793fb27b04297378178ff3d430dcdf

Observation e8217cb1-0a4a-49f7-9ed5-c678a3b68b78 · outbound

This paper cites Mme: A comprehensive evaluation bench- mark for multimodal large language models.

Object-Centric Vision Token Pruning for Vision Language Models Mme: A comprehensive evaluation bench- mark for multimodal large language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.654488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.654488Z digest=sha256:4de23f226ab8dc67fd0d9338bc8dba5d7381d9a0e98b8cf55b0e66f16bec113c

Observation 31db8585-1472-4072-b029-37f2c5229368 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Object-Centric Vision Token Pruning for Vision Language Models Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.657702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.657702Z digest=sha256:a9580fd49bc14c8da0cfba19c8f90d0a8601ecb43ffc16eb2e38907e7cfa11e3

Observation 609920c9-3a26-4a6e-a95e-ef91aa253854 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Object-Centric Vision Token Pruning for Vision Language Models Vizwiz grand challenge: Answering visual questions from blind people

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.661068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.661068Z digest=sha256:fd587d3f498b84c0bf2d8f63b154482a3ee086ff1deb9ffcc899d14b078f9598

Observation 663eec7e-fe31-49d6-a943-7b3e1687816f · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Object-Centric Vision Token Pruning for Vision Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.664289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.664289Z digest=sha256:7329b4b413f7f586070da03d58650e5a3201de25c88994364fe4263b38c7be61

Observation 54797207-daea-4a8f-ab57-2a0b46532e1f · outbound

This paper cites Improving Object-centric Learning with Query Optimization.

Object-Centric Vision Token Pruning for Vision Language Models Improving Object-centric Learning with Query Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.667395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.667395Z digest=sha256:ba1e980231c0ec8ac9dbf341e1d62d24440105e5a3dc510335f73139eb5d519a

Observation e1d3a1af-e38b-4fa0-9683-b2c3d9861204 · outbound

This paper cites Spot: Self-Training with Patch-Order Permutation for Object-Centric Learning with Autoregressive Transformers.

Object-Centric Vision Token Pruning for Vision Language Models Spot: Self-Training with Patch-Order Permutation for Object-Centric Learning with Autoregressive Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.670107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.670107Z digest=sha256:159af2116f115ee1499d63528e51d2f3253b2ffd2e6a194e205df154637d15ff

Observation 29e27677-42f9-4bd0-a46a-886eb9446861 · outbound

This paper cites Seed-bench: Bench- marking multimodal large language models.

Object-Centric Vision Token Pruning for Vision Language Models Seed-bench: Bench- marking multimodal large language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.673071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.673071Z digest=sha256:c085169b71dbf34a57aefc0038b3354bf5ac4c804115b415af064f25641840bf

Observation 9f61697f-ab6b-404a-8f56-d7c1bd16f01a · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Object-Centric Vision Token Pruning for Vision Language Models Evaluating Object Hallucination in Large Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.675843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.675843Z digest=sha256:0843eb6d14d27fd81edc32d92830959f3ae6ffeb6e0b923b1bf6238be2c0541e

Observation 8648c6b5-3eec-4c1d-ba4a-0156d6f1239a · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Object-Centric Vision Token Pruning for Vision Language Models Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.679008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.679008Z digest=sha256:92557cdcfb97115b92464ac106b1320bcc7a5a38d08bd80c0a9d1febb31e948d

Observation 781f2851-0b89-49b7-af6b-9cdee1c13a93 · outbound

This paper cites Microsoft coco: Common objects in context.

Object-Centric Vision Token Pruning for Vision Language Models Microsoft coco: Common objects in context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.682081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.682081Z digest=sha256:06b0f5c17cb4d2125a95395a66fce5eb6d5d70dd9ef9edfee9b75ee5c2062284

Observation fc242561-69c5-40b5-9219-65e5fd4cfde7 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Object-Centric Vision Token Pruning for Vision Language Models Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.685000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.685000Z digest=sha256:f88615e7c2a7ac886909080eeee809393fa19bfeb573531acdf9400610f4f939

Observation d675fbc9-106a-4e70-b47b-0b320e632b6d · outbound

This paper cites Llavanext: Improved reasoning, ocr, and world knowledge, 2024.

Object-Centric Vision Token Pruning for Vision Language Models Llavanext: Improved reasoning, ocr, and world knowledge, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.687748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.687748Z digest=sha256:6b9b395bc4c7f45a0e876a027cd04d13d9bfddb3dd35ce873cbaaa53f2370afd

Observation 2701cd67-e999-4c85-95a1-d10392917ddd · outbound

This paper cites HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models.

Object-Centric Vision Token Pruning for Vision Language Models HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.690392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.690392Z digest=sha256:f90b69b3ab9593e5e8b96431740092aa487100dc7364935e6551ed8bb7da25bc

Observation 15fe81ce-0ad1-4070-b11a-0052199e36e1 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vi- sion, pages 216–233.

Object-Centric Vision Token Pruning for Vision Language Models Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vi- sion, pages 216–233

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.693646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.693646Z digest=sha256:960392d70fd6626483e8f4b070af9ad209c0f4aa6cc602e3aa197a14ab118225

Observation ff3236da-a8f7-46d3-88c0-312359f14136 · outbound

This paper cites Object- centric learning with slot attention.Advances in neural in- formation processing systems, 33:11525–11538, 2020.

Object-Centric Vision Token Pruning for Vision Language Models Object- centric learning with slot attention.Advances in neural in- formation processing systems, 33:11525–11538, 2020

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.696316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.696316Z digest=sha256:e1a8c1d1f71b8ede75f23b190225cf6471e0fdc6678a37368a7c298d10540422

Observation 87a856d0-8a55-40b1-ae5f-547db0fc364b · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,.

Object-Centric Vision Token Pruning for Vision Language Models Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.699067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.699067Z digest=sha256:cc9efbea6a4eb4e25f8e17e33516b2299627c38855fcf1a5862c6c4263384bcb

Observation 44d11a48-ed1b-4212-8026-ed04f7018e8b · outbound

This paper cites Temporally consistent object-centric learning by contrasting slots.

Object-Centric Vision Token Pruning for Vision Language Models Temporally consistent object-centric learning by contrasting slots

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.702079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.702079Z digest=sha256:c286714d8890479bf58f8ba27a95a7fecb20d747ea18acb8b271ee97cf0318e9

Observation cfb0f784-a140-4f1c-9729-15bada9374dd · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Object-Centric Vision Token Pruning for Vision Language Models Learning transferable visual models from natural language supervi- sion

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.704723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.704723Z digest=sha256:2f34a8ba78f4db6dc746ea5e17ff10caf94f88e927c705fd90b4f301ef69c325

Observation a99ef534-51d8-492f-af74-ebadcf5362c6 · outbound

This paper cites Bridging the gap to real-world object-centric learning.

Object-Centric Vision Token Pruning for Vision Language Models Bridging the gap to real-world object-centric learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.708009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.708009Z digest=sha256:c2fce9247d7d13e2ccacd50a3619f1fe4bf6c02666b3aa3e7ee41c0403093c61

Observation bd5e67b3-50e3-4a13-bdc5-4586ec7cc3b9 · outbound

This paper cites Towards vqa models that can read.

Object-Centric Vision Token Pruning for Vision Language Models Towards vqa models that can read

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.711377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.711377Z digest=sha256:259c49b470f583851558fc32bf09b98e6f0b9363145fc51a814e85f00330075e

Observation e037fa8f-5851-4b95-ad99-51695ad72b50 · outbound

This paper cites Less is more: A sim- ple yet effective token reduction method for efficient multi- modal llms.

Object-Centric Vision Token Pruning for Vision Language Models Less is more: A sim- ple yet effective token reduction method for efficient multi- modal llms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.714392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.714392Z digest=sha256:66394f61665bab90bbf6bf766d3e04d94e4f0b438a59963d9033ce594d5fee38

Observation 12614eef-ba43-45c0-a693-6ef5fa328563 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063,.

Object-Centric Vision Token Pruning for Vision Language Models Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.717592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.717592Z digest=sha256:f0418a98e9b60c04b84be448a9b19e9c883358d9d06dbac479350e0824924224

Observation 8744c380-d06b-41a7-81a0-c22e11551ea8 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Object-Centric Vision Token Pruning for Vision Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.721567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.721567Z digest=sha256:7ca86aaa21b0915869b855bfde3de7b966b4c1a30d1775ae760e8a55c5512340

Observation ad154c0f-9850-42c5-bdce-09a2c25ee2a1 · outbound

This paper cites SlotDiffusion: Object-Centric Generative Mod- eling with Diffusion Models.Advances in Neural Informa- tion Processing Systems, 36:50932–50958, 2023.

Object-Centric Vision Token Pruning for Vision Language Models SlotDiffusion: Object-Centric Generative Mod- eling with Diffusion Models.Advances in Neural Informa- tion Processing Systems, 36:50932–50958, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.727020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.727020Z digest=sha256:569d60d6392e6e4c46f301a6ebe9bb570e978a8954a018f96ffdfe5ac02ca5d3

Observation 61340193-c238-4791-8e28-0eb470e7efe9 · outbound

This paper cites Conical visual concentration for efficient large vision-language models.

Object-Centric Vision Token Pruning for Vision Language Models Conical visual concentration for efficient large vision-language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.731279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.731279Z digest=sha256:4f6a457ff87b4bc8e1471223d4709a296b7cbf48c298995b35a539f9a45c829e

Observation c0f215aa-3b7b-4206-b922-b7aaca170ec1 · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models.

Object-Centric Vision Token Pruning for Vision Language Models Visionzip: Longer is better but not necessary in vision language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.734872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.734872Z digest=sha256:0f4a6b03f3340792203e89bd1884dc9446fb95c52064ce6e4acd6f5be384e2de

Observation 501c6f9c-7aad-4093-845b-4ac155d10106 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

Object-Centric Vision Token Pruning for Vision Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.738313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.738313Z digest=sha256:7e574a7f948db28e81012e4ee73623e9894d7773b88982397bf984363dab34fe

Observation 84f066d8-5d1c-4deb-a718-97079d7bd7df · outbound

This paper cites Object-Centric Learning for Real-World Videos by Pre- dicting Temporal Feature Similarities.Advances in Neural Information Processing Systems, 36, 2024.

Object-Centric Vision Token Pruning for Vision Language Models Object-Centric Learning for Real-World Videos by Pre- dicting Temporal Feature Similarities.Advances in Neural Information Processing Systems, 36, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.741594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.741594Z digest=sha256:8c2dcfaaf38c6a5608b67301cc4acbbd92b420128c660df6ece45bdb05782bee

Observation 22b0e0e0-5ed3-4dbe-863b-fbb411463446 · outbound

This paper cites Lmms-eval: Re- ality check on the evaluation of large multimodal models.

Object-Centric Vision Token Pruning for Vision Language Models Lmms-eval: Re- ality check on the evaluation of large multimodal models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.744475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.744475Z digest=sha256:17d28d52591c843f4776da05edc8a6fe7cb6124b9bd99c8e136c514d3610dd65

Observation 41fcc725-f991-4930-89b4-fe565ca1e56d · outbound

This paper cites Sparsevlm: Vi- sual token sparsification for efficient vision-language model inference.

Object-Centric Vision Token Pruning for Vision Language Models Sparsevlm: Vi- sual token sparsification for efficient vision-language model inference

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.747622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.747622Z digest=sha256:876a48f8f8b62c91c29e682effe8eaa0722a170b92f95b4628fdafe5fd9ed440

Observation e4d2496a-9011-448c-bea4-4daf8586d8d1 · outbound

This paper cites Predicting video slot attention queries from random slot-feature pairs.arXiv preprint arXiv:2508.22772, 2025.

Object-Centric Vision Token Pruning for Vision Language Models Predicting video slot attention queries from random slot-feature pairs.arXiv preprint arXiv:2508.22772, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.751167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.751167Z digest=sha256:eb173f37ef41bd2c9108cad297541ae98af79bb16b4de33825763dcb4879d081

Observation 15f17c84-6894-4ad7-a2af-1e1a4649d094 · outbound

This paper cites Vector-Quantized Vision Foundation Model for Object-Centric Learning.

Object-Centric Vision Token Pruning for Vision Language Models Vector-Quantized Vision Foundation Model for Object-Centric Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.754330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.754330Z digest=sha256:52f6d8b734aa86401f0c317670945eaebb5992caa5def96344a89d9bb43f6d30

Observation 66de32da-b792-4fa5-97a9-c7195582385b · outbound

This paper cites Smoothing Slot Attention Iterations and Recurrences.

Object-Centric Vision Token Pruning for Vision Language Models Smoothing Slot Attention Iterations and Recurrences

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.757841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.757841Z digest=sha256:450472b41a86e71946669a385c9faee5906fb1f7adb99a5025239d265cc62b8c

Observation cc776f21-cf19-4a1e-a9fb-260e27c90d8d · outbound

This paper cites Slot Attention with Re-Initialization and Self-Distillation.

Object-Centric Vision Token Pruning for Vision Language Models Slot Attention with Re-Initialization and Self-Distillation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.761323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.761323Z digest=sha256:95f46b9b8f16354ee334beb439a2f06de65259962b6c488678351ab9cb51674c

Observation 8dea26c4-5443-416e-9183-a419fbf7c9e1 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Object-Centric Vision Token Pruning for Vision Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T20:19:57.764281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:19:57.764281Z digest=sha256:7b8b2b323a1beb1ee2878886766d4182c34ec4a0c5706900811f2a1235f661f9

Pith citing papers

No inbound Pith citation observations are available.