Pith. sign in

Paper Citation Record · LEDGER

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads

As of 18 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 0 inbound Pith citation observations for arXiv:2506.03433.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03433 v2

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:09:33.397262Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

96 of 96 outbound references displayed

  • verified exact0
  • verified fuzzy61
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3259072f-c0c8-465c-8dad-8d1ea6d68f7c · outbound

This paper cites Attention attention everywhere: Monocular depth prediction with skip attention.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Attention attention everywhere: Monocular depth prediction with skip attention

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:25.503220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:25.503220Z digest=sha256:4ea8f286db60a83533d448a2752cb0dd7da5768556401248db374ca0a990e536

Observation 5a6dff74-e5a1-4030-bc6e-f33e08baecd7 · outbound

This paper cites Self-supervised learning from images with a joint-embedding predictive architecture.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Self-supervised learning from images with a joint-embedding predictive architecture

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:25.562474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:25.562474Z digest=sha256:4318b2534b488b38820b9397ac3ed213938b76eeaf3f32c527bd96bc8d4fdccc

Observation e73c98c1-debc-4e24-b8ee-4e211477e1ae · outbound

This paper cites Foundational Models Defining a New Era in Vision: A Survey and Outlook.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Foundational Models Defining a New Era in Vision: A Survey and Outlook

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:25.650110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:25.650110Z digest=sha256:0a535769c570993f259b8764632df4d4a7e2682e78dab537031bfb55e1e174bb

Observation 6ba129c2-f5bd-4a00-9077-29881ed2d33c · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:25.727381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:25.727381Z digest=sha256:c0f3bcb74c6c09dc19ea1c20682c5ff8dc85898c8b3d381e89989a1edf75ab6e

Observation 2b5833a3-8193-4182-a5db-69f97e246698 · outbound

This paper cites Beit: Bert pre-training of image transformers.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Beit: Bert pre-training of image transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:25.804823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:25.804823Z digest=sha256:4dd6de126c37393307d0cb0ac9d576772dbaceb6a4b22cab0c08c030b8e0cbc6

Observation 7dc0cbb8-fa7a-487f-899f-9edc42513260 · outbound

This paper cites Adabins: Depth estimation using adaptive bins.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Adabins: Depth estimation using adaptive bins

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:25.877859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:25.877859Z digest=sha256:245d47aa8a8246da19c6bcaeb2deb0030dde068b0dc145f404ce9c4ec843dbb1

Observation 015458de-a05f-44a2-b47e-cb13db6d2f52 · outbound

This paper cites Language Models are Few-Shot Learners.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Language Models are Few-Shot Learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:25.951784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:25.951784Z digest=sha256:ce95444f51d67c265d45f71f0d14053d7042376bc8b584def46dd175c138ad58

Observation c79ea6be-af6d-4428-ae06-59686630daef · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Emerg- ing properties in self-supervised vision transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:26.014027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:26.014027Z digest=sha256:503b2b4de6d205b1c856a71b6a81871e4e2027c23783113b761753d804771548

Observation f1cfa443-a58f-4865-9e6e-c8e2b946b5b0 · outbound

This paper cites MMDetection: Open MMLab Detection Toolbox and Benchmark.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads MMDetection: Open MMLab Detection Toolbox and Benchmark

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:26.091367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:26.091367Z digest=sha256:e06facc7bc6a3ffe5e28e5db0b76b3b0b86b538c85763685d65659cecb89a60b

Observation 7ae5764e-d0bb-4e12-88dd-7b3ebd393ebb · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:26.157579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:26.157579Z digest=sha256:a45688cef6ae8eba0657bafb88d476b66cbd1c94dfec2ed7e6b6c03cce7224ab

Observation 41a51fa6-5d03-4d66-a720-74aeed9eb11b · outbound

This paper cites Mixformer: Mixing features across windows and dimensions.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Mixformer: Mixing features across windows and dimensions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:26.216007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:26.216007Z digest=sha256:021c71e6ae773df68d07f79046eb43ea0358752e34051d2394c9dff6a724a848

Observation e1627e6d-8d87-4f4f-b0a3-da425aa2dc82 · outbound

This paper cites Adaptformer: Adapt- ing vision transformers for scalable visual recognition.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Adaptformer: Adapt- ing vision transformers for scalable visual recognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:26.304768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:26.304768Z digest=sha256:35d24a08fdf4858d09a3af4169b24fb144f3a69a182a0ca18acf5903b600382c

Observation f29d0633-c3aa-42a1-b7ef-8b85eda17217 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads A simple framework for contrastive learning of visual representations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:26.365979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:26.365979Z digest=sha256:3164910b1bcdf6b1546a553e3cc37b8d1105cc01e65bff608f0c8fdb34198d05

Observation 8a9c32a1-630e-43b9-80ce-35eefb827812 · outbound

This paper cites Vision transformer adapter for dense predictions.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Vision transformer adapter for dense predictions

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:26.435734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:26.435734Z digest=sha256:278d6dca2597c296a0d4fc1d59a01505f4a2876c8e31751d7362dae891db7fea

Observation 3ec0a4e6-e918-451e-907c-fc9cf16c620e · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:26.502832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:26.502832Z digest=sha256:cf7f43fab58d7fd608f62127da5a402af2832da130eb8bcea67bfe3adb71f02c

Observation 4b000884-592d-44fe-8e69-dfa7b8495789 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Masked-attention mask transformer for universal image segmentation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:26.571580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:26.571580Z digest=sha256:11cf10958501761a8136836e186aa5128475a19013bdfd6048afa0ff4cef5b7a

Observation be2d17de-c24f-42a2-8a14-b95afb121481 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:26.716702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:26.716702Z digest=sha256:d210f98a8f053a2ae93733e7efaa34435e15317cda37729cf7219a0344aeba93

Observation 8e461f47-cf2c-4408-8883-c89bdcf751d9 · outbound

This paper cites Twins: Revisiting the design of spatial attention in vision transformers.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Twins: Revisiting the design of spatial attention in vision transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:26.893912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:26.893912Z digest=sha256:a1e20eb03108e7db475e6edbe80df05c6ca9fb135c8f8f7c9403718065c2a98f

Observation 22d2b970-7161-4f76-a774-3ebcb701146b · outbound

This paper cites Mmsegmentation: Open- mmlab semantic segmentation toolbox and benchmark,.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Mmsegmentation: Open- mmlab semantic segmentation toolbox and benchmark,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:26.953108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:26.953108Z digest=sha256:e2c7594fd6d2686a6cb55d3475f617708223beea4d1d505d35db4fce7f65f25e

Observation 89ba021c-d6f1-44de-9e27-7679cb9af83c · outbound

This paper cites The cityscapes dataset for semantic urban scene understanding.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads The cityscapes dataset for semantic urban scene understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:27.156861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:27.156861Z digest=sha256:c23ce0f6011805d53e7550e57af258d0c6c0729d1a47f371077cd1f9698b34b4

Observation f1dfef07-8d21-4eb9-89d1-5487746f385a · outbound

This paper cites Deformable convolutional networks.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Deformable convolutional networks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:27.322128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:27.322128Z digest=sha256:543b55baf3eef06b3245cb4592ba12202fec66e2fadebcc9332a87ab60ae956f

Observation 97a62938-dd01-4df9-9635-99b541c44670 · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Instructblip: Towards general- purpose vision-language models with instruction tuning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:45.255268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:27.472002Z digest=sha256:41a0b57455017683bf9924468abc4e5d4b5607a22c0305f1fcdf2815e2b50eeb

Observation b61b55ad-462f-47c2-84d7-a8c558f6347e · outbound

This paper cites Scaling vision transformers to 22 billion pa- rameters.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Scaling vision transformers to 22 billion pa- rameters

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:45.180387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:27.616858Z digest=sha256:7316d3f92d9e0312cb4ec3ec52d2e4ae56a0f246cdb8bf3f5bd28d6e4d5bd925

Observation 7d531c04-b57e-41e6-bd97-e66b886870ac · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads An image is worth 16x16 words: Trans- formers for image recognition at scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:27.826601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:27.826601Z digest=sha256:78fc19c2f4c3a0697271f7a991760c94ed070166cdaa34dedd57eb12d10b6027

Observation e2fb3560-fa9d-430a-9bf8-7d30682d6726 · outbound

This paper cites Eva: Exploring the limits of masked visual representa- tion learning at scale.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Eva: Exploring the limits of masked visual representa- tion learning at scale

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:44.959574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:28.059819Z digest=sha256:099bc17dfcc95b5c93156b789b731f83ed168cf85e547511926cb3e88d828cd6

Observation e99d77dd-cc26-4d5e-8413-36dbbd58fb33 · outbound

This paper cites Deep ordinal regression net- work for monocular depth estimation.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Deep ordinal regression net- work for monocular depth estimation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:44.652414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:28.181241Z digest=sha256:bbd14186a0d1164d83777550785c1209972a07f7130621afda0885cf0bc37e19

Observation 3fc27194-76cc-4bc4-ba45-1e646738c68a · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:44.431782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:28.307437Z digest=sha256:657f912466ee658db412d7aed5bf1165c4f31b3b6d1c52225eed002876d45d05

Observation a9e5e659-6a19-4b56-947b-115dfb8a7e96 · outbound

This paper cites Bootstrap your own latent-a new approach to self-supervised learning.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Bootstrap your own latent-a new approach to self-supervised learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:28.417640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:28.417640Z digest=sha256:9af548d5638e3cc44124424866db94d571f5a9c027fe3dc6e4c321c32ec30a6e

Observation 45015805-4052-4acf-9c11-d527202d93be · outbound

This paper cites A survey on self-supervised learning: Algorithms, applications, and future trends.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads A survey on self-supervised learning: Algorithms, applications, and future trends

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:44.169139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:28.529369Z digest=sha256:68ade2611898c2856c29a7c148f857344f199ff52a8450f978962317d861bef1

Observation f47fa069-4e62-48af-a500-1b9cf28921cd · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Vizwiz grand challenge: Answering visual questions from blind people

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:44.132535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:28.660175Z digest=sha256:1f03a863d12e05c3702517f8cb3d6c26dd1a3bb11cb66d0cb1d85bde7b04d8dc

Observation 7809092b-2956-4525-a06b-6afbddc61f86 · outbound

This paper cites Flatten transformer: Vision transformer using fo- cused linear attention.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Flatten transformer: Vision transformer using fo- cused linear attention

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:43.979734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:28.781953Z digest=sha256:1d7682314c44ffa83a36d7fa284da7e0895b968bbacb68a8eba37ab74fef42f6

Observation 5edb397e-3521-460d-8699-efba01735197 · outbound

This paper cites Deep residual learning for image recognition.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Deep residual learning for image recognition

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:28.900404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:28.900404Z digest=sha256:ca6e08d29f1aa6907c31297776e1c32b330863809e761e5e84b0f19615e5f322

Observation efb7723a-88fc-4cd1-965a-18ab8b0b525c · outbound

This paper cites Mask r-cnn.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Mask r-cnn

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:43.851888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:29.007919Z digest=sha256:032191eed013e57a4981e4a653f45f40eebb1580b29043847770d902437eedb4

Observation 2983936a-1c41-4978-9176-afedc02cfe0d · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Momentum contrast for unsupervised visual rep- resentation learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:43.704596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:29.127166Z digest=sha256:6437e5dfef845d8e662f9be0388f9eac81d9abd0e7755d3c062703e4fbd32971

Observation 1bbbe33b-d901-4c99-aee8-ec89037e9312 · outbound

This paper cites Masked autoencoders are scalable vision learners.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Masked autoencoders are scalable vision learners

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:43.450809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:29.277013Z digest=sha256:e8f09fad61c31277a76a7e2cf08ccf86f586dd0adf7d75f72bf969d21d64fb07

Observation fd28fc32-2224-4c36-85e3-d8335af63ef9 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Lora: Low-rank adaptation of large language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:43.236551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:29.400922Z digest=sha256:62f66964847809e8ccd444a63d8f5809e9cee230b02afd24756a8a08f45bed2c

Observation b2a137e0-97db-425e-84c9-ef78873510ca · outbound

This paper cites Introducing idefics: An open reproduction of state-of-the-art visual language model.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Introducing idefics: An open reproduction of state-of-the-art visual language model

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:43.026866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:29.511917Z digest=sha256:711a814041ff5073c27f8d6dcab1eb478018d884863f0f95c10441b32015fde8

Observation 442c063a-cfd3-441f-abc8-4d0d5d5b8b4b · outbound

This paper cites Oneformer: One transformer to rule universal image segmentation.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Oneformer: One transformer to rule universal image segmentation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:42.770096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:29.629874Z digest=sha256:18811968b785647e7d0a1baae00ae2e4a01e46db7dfd616cc4978ab32b828e9e

Observation 4b8870aa-0c84-4aab-8234-897454c6b332 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Scaling up visual and vision-language representation learning with noisy text supervision

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:42.424384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:29.727886Z digest=sha256:7b63f9e0cb7ec6b0de449f4c8936eb0d4b0b8e18ac339604e4d5822da172bef8

Observation 38f5c9eb-7431-4405-9d1b-6e7d0b2c5bdd · outbound

This paper cites Vi- sual prompt tuning.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Vi- sual prompt tuning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:42.116748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:29.759347Z digest=sha256:459709dd3485bd88a2d58f768e2abeb20321920dd66ac657b3ce6d670d474cdc

Observation 967c7d6c-dad3-4300-b192-e839dd0a1f87 · outbound

This paper cites Convolutional Bypasses Are Better Vision Transformer Adapters.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Convolutional Bypasses Are Better Vision Transformer Adapters

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:29.789945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:29.789945Z digest=sha256:bbca096f584400f0ea37ee46e7c48cb29feb7ddf4cf6bfd6e6f6ee70d792ab16

Observation 530a8c4c-5fee-4b4a-b166-753686497d3c · outbound

This paper cites Fact: Factor-tuning for lightweight adaptation on vision transformer.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Fact: Factor-tuning for lightweight adaptation on vision transformer

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:41.890658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:29.815986Z digest=sha256:d63444d5fbac5185cb906a3adf3289462c382550c93f23a3415208ecc3fadbbb

Observation 6824f71b-b31d-459c-912f-fdc7ad2c62c4 · outbound

This paper cites Segment any- thing.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Segment any- thing

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:41.606310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:29.847153Z digest=sha256:16ba4ce59205053f0bd9c57e99a013be996d5fd3b011b4d498752bc1faeb4b31

Observation 38eac8ca-5c8a-400b-bc4a-c0c6258cc369 · outbound

This paper cites Similarity of neural network representa- tions revisited.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Similarity of neural network representa- tions revisited

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:41.308185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:29.877229Z digest=sha256:206f928d7350db128bb1aa0416fb6b9c69be457dbb1ff0dfeb168d8723ff7b80

Observation 7ea0415c-00f6-4ded-b754-5d323829846a · outbound

This paper cites Mask dino: Towards a unified transformer-based framework for object detection and segmentation.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Mask dino: Towards a unified transformer-based framework for object detection and segmentation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:41.037793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:29.910879Z digest=sha256:33f83429158f119a88baee54a83ce0cc0f956537a251f2c5e07eaf59649ca14d

Observation 940846d3-e1ca-4e71-87c3-9ec17420200d · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:40.786115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:29.954039Z digest=sha256:ae02632ada4d2987ad31c9dfb00e5cd30156d78b7d008fb97d930367284498c6

Observation 23514c34-1a8e-4d4a-9c7e-d94883a6abcd · outbound

This paper cites Benchmarking Detection Transfer Learning with Vision Transformers.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Benchmarking Detection Transfer Learning with Vision Transformers

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:29.983561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:29.983561Z digest=sha256:d4881142b788e838fa845e8989f60505a951233172cb57d4ff683c6939c3d8d5

Observation 32c03cab-0be8-4dfe-a5bf-80fc4e7dc4f2 · outbound

This paper cites Exploring plain vision transformer backbones for object de- tection.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Exploring plain vision transformer backbones for object de- tection

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:40.500008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:30.020550Z digest=sha256:de999aed9d224e4f590d1e0c3c0c1ccd7dbd9241d1501251e58168275086392c

Observation 3c753889-975f-4647-b694-b46ae19bb292 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Evaluating Object Hallucination in Large Vision-Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:30.053409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:30.053409Z digest=sha256:b7e10b2542d7832fc2a728d3ba24fa8b0642c0e0be4e61162d211894569b66bb

Observation f2bfbacc-3d5b-48d5-a79f-f8303ca9cd2c · outbound

This paper cites Visual Large Language Models for Generalized and Specialized Applications.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Visual Large Language Models for Generalized and Specialized Applications

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:30.084456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:30.084456Z digest=sha256:259e2959ee0cdef889766e6d1b78c576358f027bc77b2e1742d401c7297b567e

Observation aedae50e-d8c5-429c-9a71-1036f9d57010 · outbound

This paper cites Binsformer: Revisiting adaptive bins for monocular depth estimation.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Binsformer: Revisiting adaptive bins for monocular depth estimation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:40.254830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:30.125615Z digest=sha256:0955ff5617bba625db6c4b22cfdcf5df4cf1cb873c30b7e05deaf0ef90264428

Observation 221cc480-74af-484c-a42a-739ca599f18c · outbound

This paper cites Microsoft coco: Common objects in context.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Microsoft coco: Common objects in context

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:40.127811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:30.174830Z digest=sha256:9a9799aacf211c58cc151887e24754684d184f4d3c2b887dedc169a35dfc690d

Observation 402d1373-bcb3-4db1-be39-1f37bdff4500 · outbound

This paper cites Va-depthnet: A variational approach to single image depth prediction.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Va-depthnet: A variational approach to single image depth prediction

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:40.016625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:30.212604Z digest=sha256:d92f3487884a5395271415fbd9d7cb1f0bf4a5e31e76c06f775568373af4493a

Observation b768bfb2-ceee-42df-8090-c5d879a4a74d · outbound

This paper cites Improved baselines with visual instruction tuning.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Improved baselines with visual instruction tuning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:39.840043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:30.265002Z digest=sha256:e218787d8abc237d33d295c1ac42db26755b2b219dc8ecd774feff566aa3acdc

Observation 6edd2ff8-1d5b-47dd-a504-0b6a6778b6a7 · outbound

This paper cites Visual instruction tuning.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Visual instruction tuning

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:39.673372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:30.308517Z digest=sha256:342ee01e2083f177b365648206ceefb2f9e2dde0c4abd79a9c86315b06b576a8

Observation 1e7808ca-2b1e-4969-aa9b-c350a195dae5 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233, 2024.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Mmbench: Is your multi-modal model an all-around player? In ECCV, pages 216–233, 2024

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:39.540123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:30.345111Z digest=sha256:9eadae5315fcfa3c266cbf8494822ba4734c22aeea2517c4f71c8ac96234f0eb

Observation 1892988f-deeb-4a54-b3c8-3f6bf66b9eb5 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Swin transformer: Hierarchical vision transformer using shifted windows

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:39.389456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:30.380457Z digest=sha256:65a707d48a60e1107f004dfaa70ccdc750a0c21d6a763a22ed980e4ae77d292a

Observation d909d31e-cc09-41ba-aa0e-caadb06e847b · outbound

This paper cites Swin transformer v2: Scaling up capacity and resolution.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Swin transformer v2: Scaling up capacity and resolution

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:39.235168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:30.420759Z digest=sha256:9648a49e8662a587826d9ee6c3e79758be0baca14ee866b820d4b781bdcba797

Observation 6ed1515f-7dfb-427c-ab4c-d713d70df107 · outbound

This paper cites A convnet for the 2020s.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads A convnet for the 2020s

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:30.465268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:30.465268Z digest=sha256:5dbc0968e59aa36a0f14004ffbf3795ca28267fc5d2f8ddb8367731cf31628bf

Observation 4ce4e11f-e25b-4838-a428-f1eabc752a1f · outbound

This paper cites Decoupled Weight Decay Regularization.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Decoupled Weight Decay Regularization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:30.511672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:30.511672Z digest=sha256:0fec322ce9ac7b88d850323b447575039186eb942173af59708d6daf7b925873

Observation 6311134e-0259-4838-a3c6-dba2c3df9297 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:39.109290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:30.560640Z digest=sha256:c9d3c095cd2c278fdf0566a251d83f4c21be3d55854591de59ec7461f07866e6

Observation 1e8725e5-96fc-4b70-a576-2f353482bd17 · outbound

This paper cites Time-memory-and parameter-efficient visual adaptation.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Time-memory-and parameter-efficient visual adaptation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:38.957515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:30.622635Z digest=sha256:6f52b6ce44e83d8c4f1232a8db6fd08e2d3b2e10cc7d6fca0d44f3fdbe4586c4

Observation 3da4c4c1-c22e-4f7e-bf3f-9712bfd7926e · outbound

This paper cites The role of context for object detection and se- mantic segmentation in the wild.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads The role of context for object detection and se- mantic segmentation in the wild

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:30.668483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:30.668483Z digest=sha256:02977486b0a4c3fae5e8d189c224f1026da6673c1f04af80b2a3f4549f6314e3

Observation 768dd031-b549-405a-821a-9093e571fb85 · outbound

This paper cites All in tokens: Uni- fying output space of visual tasks via soft token.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads All in tokens: Uni- fying output space of visual tasks via soft token

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:38.821616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:30.720359Z digest=sha256:00fe5a262c22a3bab0828e512020e026d65c0c7f39e51b58e04862ad85fa934a

Observation a8e34fd1-ed0c-4efd-9b2b-ac77a0aaffaf · outbound

This paper cites Dinov2: Learning robust visual features without super- vision.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Dinov2: Learning robust visual features without super- vision

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:38.655320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:30.783766Z digest=sha256:425d0d3d8cfe09d1720a811246eddeea85f8db9a00f1f4c5602d372571b5ba20

Observation 1d470de2-7de9-4c78-931c-8d2c1eddd2f3 · outbound

This paper cites St-adapter: Parameter-efficient image-to-video transfer learning.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads St-adapter: Parameter-efficient image-to-video transfer learning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:38.525231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:30.864905Z digest=sha256:efdb1b6645a60c2faa67494ac463ecbc6c6cdbf5ce2369f3e0dba29b3f42c004

Observation eaaad0c8-d6bf-4eb8-8eb3-fd02189d1822 · outbound

This paper cites P3depth: Monocular depth estimation with a piecewise planarity prior.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads P3depth: Monocular depth estimation with a piecewise planarity prior

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:38.412615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:30.899319Z digest=sha256:c4a7ff8e106e690acd7144898705ab33f20c9ebe67700c0d3a88857fdafe96b1

Observation 7f985a76-b600-400d-92e2-203cf90d4391 · outbound

This paper cites idisc: Internal discretization for monocular depth estimation.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads idisc: Internal discretization for monocular depth estimation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:38.273501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:30.934606Z digest=sha256:0b3923ee3351f30eb2087c45ed5dc3a7a2135437f280633d2f8a6d0f6c170ef0

Observation 079de25f-fd81-44e1-b743-8dbb0879aa80 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Learn- ing transferable visual models from natural language super- vision

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:38.084748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:30.973431Z digest=sha256:08805d147a46350fb5866f3d9fc08148a8ea7d3ed26c4f35fc16935bf9417339

Observation c5e79113-f7c9-476e-9772-452b273caa55 · outbound

This paper cites Do vision trans- formers see like convolutional neural networks? NeurIPS, 34:12116–12128, 2021.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Do vision trans- formers see like convolutional neural networks? NeurIPS, 34:12116–12128, 2021

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:37.867793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:31.018949Z digest=sha256:619a3e5606e380f81aa6f2c26407ef51ead4fe5b31a5ef5d6a165158f7b11a17

Observation 34c1f187-4247-4add-8273-8f50d6476cfa · outbound

This paper cites Vi- sion transformers for dense prediction.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Vi- sion transformers for dense prediction

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:37.700892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:31.062085Z digest=sha256:88312d67334993b489a00775765d8b8c98e2d46d8605e4fb18b675652c783cf9

Observation 0f13137d-2af3-4c93-a9ce-fde997b11f10 · outbound

This paper cites Learning multiple visual domains with residual adapters.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Learning multiple visual domains with residual adapters

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:37.533889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:31.098133Z digest=sha256:a0630f03f8fb5afbb733dad6dd12e5c44545cee71b138e28f774180a413eb7c0

Observation 9e96e077-9e8c-4457-8363-09a34de5ffa2 · outbound

This paper cites Iebins: Iterative elastic bins for monocular depth estimation.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Iebins: Iterative elastic bins for monocular depth estimation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:37.365669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:31.179733Z digest=sha256:12266e436431dd2f42484499a6fa456fa28d04f462e80a89f0f54a83de7333d5

Observation a70ce1e0-d841-45de-94e2-4d38e66e0718 · outbound

This paper cites Indoor segmentation and support inference from rgbd images.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Indoor segmentation and support inference from rgbd images

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:37.186757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:31.259767Z digest=sha256:126a44e2aaf5021080743411843080223a97673e39817530f3548d582da837ec

Observation 8b5ff976-2a4b-4782-bbcd-1614c3ba9216 · outbound

This paper cites Efficientnet: Rethinking model scaling for convolutional neural networks.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Efficientnet: Rethinking model scaling for convolutional neural networks

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:37.004730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:31.340193Z digest=sha256:2e9c67a634b82b4129fce2b99ac32a079378af7c1614873f6821cd4e425ff84a

Observation 2ae3adb6-4b20-497d-9bf6-907e3b866b5e · outbound

This paper cites Training data-efficient image transformers & distillation through at- tention.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Training data-efficient image transformers & distillation through at- tention

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:36.824104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:31.410523Z digest=sha256:47b163fb1ed49ba746851d71d6c10cd7ecc2a90419dac19bc1b90e2881c7829e

Observation ebe46b12-e275-45aa-aa67-102b073baf57 · outbound

This paper cites Pvt v2: Improved baselines with pyramid vision transformer.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Pvt v2: Improved baselines with pyramid vision transformer

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:36.645650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:31.499091Z digest=sha256:142b7d3b3f0961f319898fa22bc001905c04f5df9f371040a3c54af21d353bc8

Observation 0c88edec-f5a5-4c64-b025-35746b277c91 · outbound

This paper cites Internimage: Exploring large-scale vi- sion foundation models with deformable convolutions.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Internimage: Exploring large-scale vi- sion foundation models with deformable convolutions

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:36.458608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:31.576541Z digest=sha256:4c56eab750b427f0bc153ccad584b43c8157f4b465db571b7bc1389528087787

Observation 287062d4-a3bf-4e7c-a07b-b40b8cf97fff · outbound

This paper cites Vit-comer: Vision transformer with convolu- tional multi-scale feature interaction for dense predictions.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Vit-comer: Vision transformer with convolu- tional multi-scale feature interaction for dense predictions

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:36.291060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:31.654238Z digest=sha256:3871abdd7d4b87b5e1904d6a3dbf232f9518df6755571a5df396fe2b3888ee23

Observation 63817e5e-53f0-4dbe-8191-3fba7dda9178 · outbound

This paper cites Unified perceptual parsing for scene understand- ing.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Unified perceptual parsing for scene understand- ing

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:36.120647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:31.743058Z digest=sha256:75bc2c4defe62c8d054f8cc7fc814ef5fce996eb2a9d7cbad722a81be86cbfdd

Observation 8c4a6ed8-c6a8-4050-b989-c8a690177d4d · outbound

This paper cites Focal Self-attention for Local-Global Interactions in Vision Transformers.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Focal Self-attention for Local-Global Interactions in Vision Transformers

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:31.821493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:31.821493Z digest=sha256:f4a0e58aa0158725d2987a9b0ac23a01f921a55ff43b1d27d8236466f1965f79

Observation e0ab7dbb-5722-43e5-b715-e735e7f345e9 · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Depth anything: Unleashing the power of large-scale unlabeled data

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:35.912858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:31.902326Z digest=sha256:b4ec21342d97e4ac0bdcbedcee3669fe73b9c964f9fcc27baee40520a4773b9c

Observation 3c4781a8-e821-41d6-b17d-c6366900841e · outbound

This paper cites Visual tuning.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Visual tuning

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:35.783031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:31.984771Z digest=sha256:021194442dfc3d36a446fa4f4a7ce40e37abc8f188e7f2a77ac786ff4c98c123

Observation ee55490a-26a2-4b93-b9fa-1deade12af1d · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Coca: Contrastive captioners are image-text foundation models

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:35.629902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:32.077506Z digest=sha256:975d6001f633e114a974cc2c602f68f8bbb1adc5e66c21f36f8a3aa5a4bfa6b4

Observation d3f3e392-39d7-424b-ac60-5e27922ff185 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:32.186257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:32.186257Z digest=sha256:d2d64454f9f158009f4ae5c0a07482cd2ccafe2ac60fc3a57dc20b323774ae6c

Observation 152b1732-5b13-47fc-9073-523dd585f361 · outbound

This paper cites Neural window fully-connected crfs for monocular depth estimation.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Neural window fully-connected crfs for monocular depth estimation

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:35.462006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:32.296037Z digest=sha256:b2bc62eae2ed2d8283bc70d91fc192bd717e17de8b4b28b66a65b171efbf17ee

Observation e91391c4-2d1e-4b2f-8e56-80ab8e863f47 · outbound

This paper cites Spanet: Frequency-balancing token mixer using spectral pooling aggregation modulation.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Spanet: Frequency-balancing token mixer using spectral pooling aggregation modulation

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:35.307375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:32.383399Z digest=sha256:5f460993baca56e4e15468f51ed7230562c278c755568b0a52b32e7f144c840e

Observation 722a9709-a5b9-4c82-934d-07082b6f963f · outbound

This paper cites Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:35.050142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:32.514627Z digest=sha256:87b86c12faea5060dc3721c781a74f366e3e7ca3901662a4582d9045c329b2da

Observation a97d7428-4760-4ec0-ac4e-0d40505cc666 · outbound

This paper cites Sigmoid loss for language image pre-training.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Sigmoid loss for language image pre-training

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:34.856362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:32.661067Z digest=sha256:80d1bdac3a704c961f564c1c01819011301db228a91aa379c764910adf7bc4d0

Observation 5f21ecd9-32a9-4f1f-8eda-75b0b2becc75 · outbound

This paper cites Memory efficient transformer adapter for dense pre- dictions.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Memory efficient transformer adapter for dense pre- dictions

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:34.680805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:32.747114Z digest=sha256:a5031e0894a82194fee0e21e7681f7ae74e833f205ac6a7760a6cabf93c5abd2

Observation 9add61c0-1e67-4eb3-9aac-443fa919854b · outbound

This paper cites A Survey of Large Language Models.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads A Survey of Large Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:32.845078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:32.845078Z digest=sha256:e1a40c8ad03a64c3fedd009733865c5296e427364e2feb1a496ed06a40c0db36

Observation b2d4daee-3c0d-4ae2-9304-95c7e4902b2a · outbound

This paper cites Semantic under- standing of scenes through the ade20k dataset.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Semantic under- standing of scenes through the ade20k dataset

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:34.418231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:32.991960Z digest=sha256:ad1fcc454c0b33756c9d30f3830ce2c773d05c78b06fb9920cd8f89652f8020b

Observation 63af6f46-6ab1-4d00-8c30-9cb4d3a1d2cf · outbound

This paper cites A Comprehensive Survey on Pretrained Foundation Models: A History from BERT to ChatGPT.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads A Comprehensive Survey on Pretrained Foundation Models: A History from BERT to ChatGPT

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:33.099450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:33.099450Z digest=sha256:80004615107947fcb978743d6b945132b04b294ce39b98df34216a27bfd39b74

Observation f6730720-017b-4f8a-8bd4-88f3cc8a04c2 · outbound

This paper cites Image bert pre-training with online tokenizer.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Image bert pre-training with online tokenizer

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:34.203844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:33.212107Z digest=sha256:a31b22f7a90a75d68efc6adc32172a79ab8c5f034c9852402ed8780451c8d8c9

Observation 8f1f3dac-08b2-47be-ac72-112bfc71bfd1 · outbound

This paper cites Conditional prompt learning for vision-language models.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Conditional prompt learning for vision-language models

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:33.984139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:33.318241Z digest=sha256:24d73ae3ddb524e5c5207dfa79b18e6307618e716fe04935b32a0eae1477f6e4

Observation d994396c-d3b9-46b3-9c55-d3ef05e317a6 · outbound

This paper cites Learning to prompt for vision-language models.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Learning to prompt for vision-language models

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:09:33.801518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:09:33.397262Z digest=sha256:3c8c827c10aa2c9033e6ed3ce5c3d5e02bff6be0ae7fe6fc36ad19e19d6d5552

Pith citing papers

No inbound Pith citation observations are available.