Pith. sign in

Paper Citation Record · LEDGER

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models

As of 31 July 2026, this Paper Citation Record lists 58 of 58 outbound references and 2 inbound Pith citation observations for arXiv:2603.14882.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.14882 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T10:31:05.743053Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-31T06:34:12.847434+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-08T14:44:32.242554Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T14:44:59.711636Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact14
  • verified fuzzy44
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c32950ae-fa91-4131-b677-39c5bb600aff · outbound

This paper cites Data augmentation with noise and blur to enhance the performance of yolo7 ob- ject detection algorithm.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Data augmentation with noise and blur to enhance the performance of yolo7 ob- ject detection algorithm

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.694263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:566698d55e5dffc99362ab5d88f34e9c6c177868132b88b6892173306658b6bf

Observation 648f6afc-2a75-4370-afd1-7968eab41f11 · outbound

This paper cites Object detection through search with a foveated visual system.PLoS com- putational biology, 13(10):e1005743.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Object detection through search with a foveated visual system.PLoS com- putational biology, 13(10):e1005743

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.731196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:0136dd13cd790f4638db10ede0a1b98e47db1c10b4e4ff10491899e636da241f

Observation 98835792-3e93-48f2-889e-9ae8bbf9883e · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.734642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:5d15c6bdcb6bc7f5e2a2c91eb30c507c5bb6c6320135386e25cf7e80b2b5221c

Observation 1dc8fe82-bf2c-473a-b38b-23e5989aab5d · outbound

This paper cites M ¨obius trans- formations revealed.Notices of the American Mathematical Society, 55(10):1226–1231.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models M ¨obius trans- formations revealed.Notices of the American Mathematical Society, 55(10):1226–1231

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.737928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:625562f4b874a5cec992f4eb9f31c48e73a0a9d7f1232349f0237a0d54fe0eff

Observation f4ae511c-044a-404e-be6e-f055bedf5d7f · outbound

This paper cites Qwen2.5-VL Technical Report.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Qwen2.5-VL Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:35:27.643584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:856d6e268b1d4093c5060d1aa93eb2f52e99cb7424fb68312e34e4f5c31b402a

Observation 0ffaac07-16ee-4144-bc56-e27e8abf039d · outbound

This paper cites A summary-statistic representation in peripheral vision ex- plains visual crowding.Journal of vision, 9(12):13–13.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models A summary-statistic representation in peripheral vision ex- plains visual crowding.Journal of vision, 9(12):13–13

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.744724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:0204269ed6136a1366e02f701b064327c97cb06e6ac250c419213c1765a48af2

Observation 1fb6bdf0-2b41-4cb2-a621-8e85202e8e35 · outbound

This paper cites Token Merging: Your ViT But Faster.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Token Merging: Your ViT But Faster

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:35:27.626152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:1b33e7910dc90db412ac14eafe843dfee25f8b0c78fbc92c33a33fbcf4e9b2f6

Observation bf716ec9-bd95-46e7-8e6d-d5420f1b131b · outbound

This paper cites A general survey on attention mechanisms in deep learning.IEEE transactions on knowledge and data engineering, 35(4):3279–3298.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models A general survey on attention mechanisms in deep learning.IEEE transactions on knowledge and data engineering, 35(4):3279–3298

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.727690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:1130a78f1dec28e8203332578e572b7600a3b4ff1fa1f9d1a2b9ea33f34523c2

Observation 27e13e15-c2fa-4f89-98ca-3ecb3745ca7b · outbound

This paper cites Convolutional neural networks develop ma- jor organizational principles of early visual cortex when en- hanced with retinal sampling.Scientific Reports, 14(1):8980.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Convolutional neural networks develop ma- jor organizational principles of early visual cortex when en- hanced with retinal sampling.Scientific Reports, 14(1):8980

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.741300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:a2bdc0d83c6ddb0bae0e1534625c2e8a1788097cafa1c1f53fd6ff97dfc69aef

Observation f9619b07-9ced-4ece-a261-43bc2bfd7a9f · outbound

This paper cites Instructblip: Towards general-purpose vision- language models with instruction tuning.Advances in neural information processing systems, 36:49250–49267.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Instructblip: Towards general-purpose vision- language models with instruction tuning.Advances in neural information processing systems, 36:49250–49267

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.672840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:af622b2b42954ff677b520840b563d409a0b77a0b8592b2c7b61f426fc6bae9e

Observation e38eead7-ab86-4145-8ce8-c03db7705e70 · outbound

This paper cites The representation of the visual field on the cerebral cortex in monkeys.The Journal of physiology, 159(2):203.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models The representation of the visual field on the cerebral cortex in monkeys.The Journal of physiology, 159(2):203

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.674871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:6344a35bfa4730c37c598e68f270fde7c6aa06d96f5a74e3e718a8b69d9ac8dc

Observation 836cf222-4930-47d3-b85f-1eeb032470f2 · outbound

This paper cites PhD thesis, Indian Institute of Technology Gandhinagar.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models PhD thesis, Indian Institute of Technology Gandhinagar

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.668644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:8393daf04ef0e08a9bb613a5c9c3e2ce4d18d359cdbbb7a71c3428caa6d70735

Observation 9bd20661-793f-4b0d-989c-b17dfc2dee5d · outbound

This paper cites Modified harris hawk optimization algorithm for mul- tilevel image thresholding.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Modified harris hawk optimization algorithm for mul- tilevel image thresholding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.747620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:17295727dcb1a6e8771fe162a08db00dfda0245db4a71b3e8b67394b9230cd45

Observation 29105450-b676-4f6f-bd49-ff593144082d · outbound

This paper cites Scribgen: generating scribble art through meta- heuristics.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Scribgen: generating scribble art through meta- heuristics

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.666373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:5c89df8bee9acee80c6514ed4b70b6114517c2752f7e5f9a1897766a9eac1e6b

Observation f520765a-c60f-4bcf-9cfd-94736ae05b6e · outbound

This paper cites Emergent Properties of Foveated Perceptual Systems.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Emergent Properties of Foveated Perceptual Systems

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:35:27.617377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:4725c7760f3c7ba4e1d4b90b68e6835866eebcec4f2b495c8190f1daca07504e

Observation 4d725512-2d7f-40e7-a844-d7a0cf57d081 · outbound

This paper cites Image quality assessment: Unifying structure and texture similarity.IEEE transactions on pattern analysis and ma- chine intelligence, 44(5):2567–2581.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Image quality assessment: Unifying structure and texture similarity.IEEE transactions on pattern analysis and ma- chine intelligence, 44(5):2567–2581

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.676882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:e8d6762a411be6840751a4e226b6b8eb656c695a0ceaf82c14b4ba035205136f

Observation 5413c99f-e089-4148-b28e-19f867774949 · outbound

This paper cites Selective attention and the organization of vi- sual information.Journal of experimental psychology: Gen- eral, 113(4):501.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Selective attention and the organization of vi- sual information.Journal of experimental psychology: Gen- eral, 113(4):501

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.681103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:983b30f8e3256bd4fff6fb6a428c077d822c2149c21021cfd02f3c0f513d2d1e

Observation 5b736cc7-cc00-46f2-8146-f9de1e5efb70 · outbound

This paper cites Metamers of the ventral stream.Nature neuroscience, 14(9):1195–1201.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Metamers of the ventral stream.Nature neuroscience, 14(9):1195–1201

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.661486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:7500401ac3d3128cbee8193e486572e67c2d8aa977d6c87bb10dbdef858ce122

Observation c080502c-a7b0-4022-8eb0-79be1878c4b9 · outbound

This paper cites The free-energy principle: a unified brain the- ory?Nature reviews neuroscience, 11(2):127–138.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models The free-energy principle: a unified brain the- ory?Nature reviews neuroscience, 11(2):127–138

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.659117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:cdf9f7bba66839b98ad200c0e0c9abf601ddabca7e5b34855ef485cd65f47568

Observation ed7dd5ed-17e5-4800-a57c-90329ed1a442 · outbound

This paper cites FeatUp: A Model-Agnostic Framework for Features at Any Resolution.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models FeatUp: A Model-Agnostic Framework for Features at Any Resolution

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:35:27.608930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:cb7f89ea25ca9ac1456da769c8efa43675f082fc4d42e8026a97811cbb38f4dc

Observation 749041e1-ea54-4138-ba47-b955f65917e9 · outbound

This paper cites Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% perfor- mance.Visual Intelligence, 2(1):32.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Mini-internvl: a flexible-transfer pocket multi-modal model with 5% parameters and 90% perfor- mance.Visual Intelligence, 2(1):32

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.762710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:a9ba159ccc45e2a41b6122bcbb1b82a6ae1ede48bd292c107370f700c7a1f9fc

Observation 52dbf0eb-e94f-4dd0-bd72-c73bb32557bf · outbound

This paper cites Vari- able resolution improves visual question answering under a limited pixel budget.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Vari- able resolution improves visual question answering under a limited pixel budget

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.650273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:dd10e18a2c2fd758c648eace68058523f87f6fc6af516c20a4f55bdb9bd408ab

Observation 990b23ff-9a13-40d3-bd34-b84cd34ab078 · outbound

This paper cites See- ing more with less: Human-like representations in vision models.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models See- ing more with less: Human-like representations in vision models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.654698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:cacaea355b57ffc40459843ea7812eee16ecc8f66dc71f3f36c9de290f9652a6

Observation 07a056aa-f506-4a0a-9fdc-862517fda58c · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.713950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:9ac65d644a9c2fa26a1c6d3e55eaaa60b2e33eda06cd8dc7e44fbbefe9f2a991

Observation 0466e7cf-1a37-4ca0-a7a6-6f679a77bec9 · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Lvis: A dataset for large vocabulary instance segmentation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.646080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:e9943c84b5a2d063b61bb9d353f90c64d381235823e1d27cd10bc282daa90460

Observation a5c62fa2-5f0a-4ae3-b774-f519f8053f40 · outbound

This paper cites A survey on vision transformer.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models A survey on vision transformer

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.643977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:a4476c7d4a18a100118b9da7690c3699c825efd48f8643aa9e6d0a60b6875b7f

Observation 82eb9942-4941-40b6-9882-1069f6e1fec9 · outbound

This paper cites Coco-periph: bridging the gap between human and machine perception in the periphery.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Coco-periph: bridging the gap between human and machine perception in the periphery

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.648058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:a4dffcd2f577869590551692ff69735a59bbea0c6baeb8c68e29b2bf1dfcb1ca

Observation 14c681cd-a103-4d04-a332-dac975a1aea9 · outbound

This paper cites Eye movements in natural behavior.Trends in cognitive sciences, 9(4):188–194.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Eye movements in natural behavior.Trends in cognitive sciences, 9(4):188–194

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.652701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:a758a7fd1d392191d1ea551c9cf3fcf80f1634b2e8e0bdab2a5190c91ece0283

Observation 79981c56-12b0-4e38-9522-7e19687406cb · outbound

This paper cites Vi- sual prompt tuning.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Vi- sual prompt tuning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.642204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:76782d5cb061b899fdc69ac779009553b08f1f2509efbc7f05c91a2ea17b1402

Observation 929a1275-463c-42d0-9969-634425b509b0 · outbound

This paper cites Transformers in vision: A survey.ACM computing surveys (CSUR), 54(10s):1–41.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Transformers in vision: A survey.ACM computing surveys (CSUR), 54(10s):1–41

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.640412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:59fcbb1c542b0f506750ce30eff24eeb73470d9c390bc216e81566907834cc89

Observation d99382e1-c4b5-440e-8e40-64155488d60f · outbound

This paper cites Foveation in the Era of Deep Learning.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Foveation in the Era of Deep Learning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:35:27.662926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:2b7979a830a8e4a8eb189f2809bd0b72c3e8f9204e27db06c6a2b8f5fabf2eed

Observation d122a807-d21f-4a6a-bf84-9729c95af1f8 · outbound

This paper cites Oxford Uni- versity Press.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Oxford Uni- versity Press

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.670665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:fb1de585655fbaa8a09aa39587223719fcbaf5cc416b9f5dc7038f4e783e3f84

Observation 7ca8e6f7-f59b-4d3c-915e-01d66a709c4b · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:35:27.613119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:67fb39f64ede0c540806443d740382ff10c3c54a05c93cd14e01e83f3d7c7049

Observation 99321951-9ab3-4bcd-938e-2702b5060799 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:35:27.647915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:493ce54a41b0b93459af210062c147fe6b1195545741dd8acb334f03e2a53692

Observation 4cc72545-f925-4f5b-9606-645e278c9179 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.638366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:34d8ba1ff54d3b99ee248cc0d703e254cb48486a2c4109def3e197936202b6f5

Observation 5063c342-afa5-446b-9dd7-63d8c0d2ed08 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:35:27.621585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:fef6979ad4c18e47d48edecfb76a859c59a97ca769da22e4d724d37aa5e2dfba

Observation 30e190e7-0b4c-42d4-871d-2a753af1ae27 · outbound

This paper cites Bio- logically inspired deep learning model for efficient foveal- peripheral vision.Frontiers in Computational Neuroscience, 15:746204.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Bio- logically inspired deep learning model for efficient foveal- peripheral vision.Frontiers in Computational Neuroscience, 15:746204

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.656836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:54097596efed3d18ff6c3789a804ab91b8157619ca1c845a2dffc089612e4fcc

Observation 7b332d7e-6c2c-4cf3-ab75-c75b15888d8c · outbound

This paper cites Mind Meets Space: Rethinking Agentic Spatial Intelligence from a Neuroscience-inspired Perspective.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Mind Meets Space: Rethinking Agentic Spatial Intelligence from a Neuroscience-inspired Perspective

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:35:27.635083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:f334c28debf58b4e56d0fb073d655055fea8deb6c2792749e34cdceede670627

Observation a03270dc-468f-497d-95dc-4e01ce2af4d9 · outbound

This paper cites SmolVLM: Redefining small and efficient multimodal models.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models SmolVLM: Redefining small and efficient multimodal models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:35:27.604164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:4ce15082b8cbfd5a0b8ba6b511ab3127cf75b2eb41b3fe14b6784cef5ec819a3

Observation 3e1398e7-de42-4952-a76d-a94d53d1d51e · outbound

This paper cites Peripheral vision transformer.Advances in Neural Informa- tion Processing Systems, 35:32097–32111.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Peripheral vision transformer.Advances in Neural Informa- tion Processing Systems, 35:32097–32111

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.664296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:4076f0addea83d0fe481ef6adf9786c0e09cde1a21cbcd61bde4126027a8a797

Observation ff5e0932-77bb-4a0e-8bf0-8a36035ca3a8 · outbound

This paper cites The geometry of m ¨obius transformations.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models The geometry of m ¨obius transformations

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.756118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:9e6a1499dbdf47b36daf0abdd0006769c037998449fbed71f1921253d4c8b0b4

Observation 5b08aeba-4620-41e4-a457-0e53e67a8a2f · outbound

This paper cites Learning to search for and detect objects in foveal images using deep learning.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Learning to search for and detect objects in foveal images using deep learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.753523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:5b2536916d27d4f770968a741abfc73497d6d17a9ea7832891d99701d9802c4b

Observation 52f67fdc-44b3-4d3c-bb24-706a049ab5cb · outbound

This paper cites Human peripheral blur is optimal for object recognition.Vision research, 200: 108083.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Human peripheral blur is optimal for object recognition.Vision research, 200: 108083

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.710594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:05e4ee8673574a35f688d94d35f056782d0c172f765eaad1c6c4c7a06d95b2a3

Observation a94cd915-bce3-4f97-b45c-40659c023607 · outbound

This paper cites Dynamicvit: Efficient vision transformers with dynamic token sparsification.Advances in neural information processing systems, 34:13937–13949.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Dynamicvit: Efficient vision transformers with dynamic token sparsification.Advances in neural information processing systems, 34:13937–13949

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.750484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:0518ebfb7d6082829a83a02c36474574076372972df20a853623a3f22ca3627d

Observation c74410da-f002-4040-a94e-d837bdc379a8 · outbound

This paper cites Biologically inspired image sampling for electronic eye.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Biologically inspired image sampling for electronic eye

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.723571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:4bad087004552f157398d924b340e1722d79e506d69d10ee1142ee140711a718

Observation eaa78c98-22e2-47ad-9a4c-19cd4889d51e · outbound

This paper cites Robotic materials with bioinspired microstructures for high sensitivity and fast ac- tuation.Advanced Science, 13(15):e09739.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Robotic materials with bioinspired microstructures for high sensitivity and fast ac- tuation.Advanced Science, 13(15):e09739

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.717368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:c66c388c1deadd8c2f2a1b853a1e023f7c298ebf7ace3e28a929ffafc4cf4ccf

Observation 0218b050-ffb4-4e88-b015-d151fb2e6cf9 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowl- edge.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models A-okvqa: A benchmark for visual question answering using world knowl- edge

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.706937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:c7348e7c98259b34d39fc0f829057ef18af69ea5f213586e7fd9bc13e2a9a0ca

Observation d7534fa2-d532-452e-90a6-86752c1da381 · outbound

This paper cites Behind the Machine's Gaze: Neural Networks with Biologically-inspired Constraints Exhibit Human-like Visual Attention.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Behind the Machine's Gaze: Neural Networks with Biologically-inspired Constraints Exhibit Human-like Visual Attention

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:35:27.658311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:13aeefe1501d9057d1bc148271db3dd061614e542b0943badcb0d145b18b4623

Observation f3921aa6-5a35-43f5-ab29-7575a8922137 · outbound

This paper cites Multivariate stochastic approximation using a simultaneous perturbation gradient approximation.IEEE transactions on automatic control, 37(3):332–341.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Multivariate stochastic approximation using a simultaneous perturbation gradient approximation.IEEE transactions on automatic control, 37(3):332–341

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:35:28.678951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:103e77e1e9fb19fe186ab633f05e01bf36a7b91a6f67404d3d298c6e68d56f77

Observation af1f6430-0ad6-415a-82e8-961a673bfa71 · outbound

This paper cites Pe- ripheral vision and pattern recognition: A review.Journal of vision, 11(5):13–13.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Pe- ripheral vision and pattern recognition: A review.Journal of vision, 11(5):13–13

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.759239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:47e676da52685bd83d6f9e8a1a8656a95337a54ecc07d4f85ff2b5fba643fe90

Observation 41923778-ee58-4d55-a2a5-73b6a9255a3f · outbound

This paper cites Central and peripheral vision for scene recognition: A neurocomputational model- ing exploration.Journal of vision, 17(4):9–9.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Central and peripheral vision for scene recognition: A neurocomputational model- ing exploration.Journal of vision, 17(4):9–9

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.697699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:19dc99bcc6af3d0444991774ac93ec3bd6786f7924bb74c3a96c3e358636c5c7

Observation d9ea73e5-70da-456b-bb51-9e6a76e7037c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:35:27.630268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:8ef1069c8bd83d5c112bc37fac4af20f4235f9d6faf9abc8b16459bd6edd5913

Observation 88e62a00-72a6-4ce4-89da-d692cf9e9866 · outbound

This paper cites Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.Ad- vances in neural information processing systems, 33:5776– 5788.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.Ad- vances in neural information processing systems, 33:5776– 5788

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.720624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:d230ac47589d3e57b1f56e2c17b87cfec3663ab56fd1c76d2e750072b4371a68

Observation 573345b3-4ac8-47be-95dd-b4a9928691c7 · outbound

This paper cites Emulating human-like adaptive vision for efficient and flexible machine visual perception.Nature Machine In- telligence, pages 1–19.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Emulating human-like adaptive vision for efficient and flexible machine visual perception.Nature Machine In- telligence, pages 1–19

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.703051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:ad64b60bcb20bfd715852b4e9017356c0f3d1c0a02d399a7b7f34769b2b1bce0

Observation 01a0c6a9-d91f-4ce8-960d-4fdeeda6d0d8 · outbound

This paper cites Controlmllm: Training-free visual prompt learning for multimodal large language models.Advances in Neural Information Processing Systems, 37:45206–45234.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Controlmllm: Training-free visual prompt learning for multimodal large language models.Advances in Neural Information Processing Systems, 37:45206–45234

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.702339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:46da3f802072c21f7a92f83855a49889101a5783f55038ec4d1d74b3ac816db9

Observation 7cba3786-890d-45cd-ae9c-bc4f6b09d91e · outbound

This paper cites Qwen3 Technical Report.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Qwen3 Technical Report

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:35:27.639521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:ce97488a15d1d24694d750d98928de8fcf40efb510348bf1b4334e5cb3aeebbc

Observation ba856bf3-d05e-4a5b-8f09-33256ac6d2d3 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:29:05.422888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:1f1aed180bef61d918c437a3ce99e8410002fce9d53268c9a568eca8707c7b48

Observation 7e47f42b-2de2-4bdb-9986-2f98378a40c9 · outbound

This paper cites Vsi: A visual saliency-induced index for perceptual image quality assess- ment.IEEE Transactions on Image processing, 23(10): 4270–4281.

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models Vsi: A visual saliency-induced index for perceptual image quality assess- ment.IEEE Transactions on Image processing, 23(10): 4270–4281

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T10:39:57.685691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-05-15T10:31:05.743053Z digest=sha256:7941d312d3bfbcbd9fcbc2c1e16dda55d3c39bbf2c65767cad2f7540befde173

Pith citing papers

Observation bc148bdd-efd0-49e6-9c82-564e8ed3f352 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models

Reference 261

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:04:01.864944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:4236af190633e79da98ad9e440fe08c6a0771977af6fd8850ad12e176a4f55ce

Observation 2ca7f7bd-a232-4a5d-b24c-c506c11375b3 · inbound

EAGOR: Embodied Reasoning in Omni-direction cites this paper.

EAGOR: Embodied Reasoning in Omni-direction LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-08T14:44:59.712899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-08T14:44:32.242554Z digest=sha256:3a599d46b6ae1ed6114a0f79e9f71fa2b6abccb3d58c6f3c567d1fe987c507ee