Pith. sign in

Paper Citation Record · LEDGER

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs

As of 9 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2507.00754.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00754 v2

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:17:14.935499Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf57972b-3253-4d67-8cc2-8dbcf447dd24 · outbound

This paper cites Language models are few-shot learners.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Language models are few-shot learners

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:24.184422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:06.174747Z digest=sha256:ab2bf9b8be15ad4f1596c4ee7eec583bd54a19be90439ea722176f21c3ffacd4

Observation 0d832a76-bf8a-4879-b8af-89e85431c105 · outbound

This paper cites In these works, they provided litmus tests for measuring how focused the attention patterns of particular models are and how they relate to model robustness.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs In these works, they provided litmus tests for measuring how focused the attention patterns of particular models are and how they relate to model robustness

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:22.984748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:11.783929Z digest=sha256:ee8cb47825c872911120e57c49c0922d01796ee007fe9882593727febb374fb1

Observation 80bb878e-7eab-488d-bf54-fac98f9d577b · outbound

This paper cites the standard deviation of the accuracy values divided by the number of different seeds.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs the standard deviation of the accuracy values divided by the number of different seeds

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:20.966369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:12.633122Z digest=sha256:5128816c59713390ae868ecc04d6056f28d328a454fc484601bf70d8b367f5cd

Observation 7df17b61-bd53-4908-bfe4-0825ff4c8666 · outbound

This paper cites Frozen Transformers in Language Models Are Effective Visual Encoder Layers.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Frozen Transformers in Language Models Are Effective Visual Encoder Layers

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:17:16.472633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:06.952393Z digest=sha256:6eb9b725d57b009361bfe9e24150b4929b73033e980a53dec79bdf8a2b679233

Observation 8067ec4b-4bd5-4eab-8942-b273bfb5247d · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:07.140211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:07.140211Z digest=sha256:23ad9600776e7fd905f91f540c6797aa1507510631fb35f00745c2078c726af9

Observation f0432910-0c1c-432f-b169-d2edd481baa3 · outbound

This paper cites Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.054862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.054862Z digest=sha256:fb3dfccdcc5dbec268a49095bf3658a194e76a60829dd05d68b3e9ff63d6c406

Observation a6a28afd-2ec9-47d8-a024-7db7890eb3fc · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs DINOv2: Learning Robust Visual Features without Supervision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.132343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.132343Z digest=sha256:af99a16d43311261ee1cbb6408251d8061d1ee18139669baea4171b465a0dd1a

Observation f5907d37-48d9-4cc7-9e57-bc8d05756ac8 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Imagenet: A large-scale hierarchical image database

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.286198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.286198Z digest=sha256:9e07e46f76308700ddff6a08b4ae1dcea66bc8278f1551c53e84db6df643051d

Observation b3724c4e-8c25-47a4-8b1f-0025c935b7d9 · outbound

This paper cites Microsoft coco: Common objects in context.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Microsoft coco: Common objects in context

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.730823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.730823Z digest=sha256:e369fdf699dc87a9cfda875051aaf3402c490815aa090fb6f4422e070befeea6

Observation dbb7affd-7b08-4c85-9bcd-6eb3e47e0215 · outbound

This paper cites Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language Models.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.954914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.954914Z digest=sha256:7fae8c20698835aa73ebe19a645bbb590c08cfa0691cb43a35dfe46f869ff583

Observation 54bfea88-6ae9-486d-80d8-fe4d093f67bc · outbound

This paper cites Torch.manual_seed(3407) is all you need: On the influence of random seeds in deep learning architectures for computer vision.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Torch.manual_seed(3407) is all you need: On the influence of random seeds in deep learning architectures for computer vision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.069919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.069919Z digest=sha256:f102686c389645c7bd2098c85c71b138488b495f257d88c09aef5f1bc3b29176

Observation 6a846dd8-e028-4ee5-ad8d-aa0c10fafebe · outbound

This paper cites Understanding Why Neural Networks Generalize Well Through GSNR of Parameters.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Understanding Why Neural Networks Generalize Well Through GSNR of Parameters

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.234749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.234749Z digest=sha256:ef5283fcc958c94068ad049e3fb0a0d31fe56556e4ad5fbe94a7a505bc9f925a

Observation a964eac8-6510-4018-b4b1-d0208ec5c971 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Distilling the Knowledge in a Neural Network

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.364748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.364748Z digest=sha256:ed40c7c26220b4467ce094d17d5baca3ab8f0af99f8b1006a9a2fc0471e4d2b7

Observation c1b79421-ce28-4c47-b15b-bc68e211d958 · outbound

This paper cites SOLO: A Single Transformer for Scalable Vision-Language Modeling.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.614748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.614748Z digest=sha256:3fc3d16cfd61b933021165d2097e9620cc2dcc350e2e2f972478c59a5796dabb

Observation 7eb34fdb-15b0-4387-afd5-973902d3b467 · outbound

This paper cites Layer Normalization.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Layer Normalization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.764750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.764750Z digest=sha256:8e5d5296fc80ddc16af7a62bb3ab42b8850e1d86646c8bdb039b133108c33534

Observation 36fc9a21-bae4-4c9b-82ca-5261a805f6d7 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Adam: A Method for Stochastic Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:10.275076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:10.275076Z digest=sha256:329cd3294e568ad32c73323bf5974494de738f1d915b0630a819206ef0f12fdd

Observation b9ac64ec-a47d-4723-b68d-338970fc7d57 · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:10.544749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:10.544749Z digest=sha256:b60be3505c65d611f8ded8e03f04c0daba0f421587c0e8f22df97da49b75371c

Observation 330ec6c4-24c3-4029-9c70-faf41c147242 · outbound

This paper cites Deep networks with stochastic depth.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Deep networks with stochastic depth

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:23.564748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:10.764825Z digest=sha256:56fd04bd8120efb71ca8352eeec26a5d8d34b4da9f4ecf907f92283e78230e2c

Observation dde40791-2763-47ba-85c7-68ab4f603424 · outbound

This paper cites mixup: Beyond Empirical Risk Minimization.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs mixup: Beyond Empirical Risk Minimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:10.884728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:10.884728Z digest=sha256:c5d1f488122a948bf5a00c5f282badfbca94abe285c3e276aee9e6bb516c6825

Observation e882166e-8e34-4822-a2a3-3639df5af93f · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:11.024571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:11.024571Z digest=sha256:6117fb84a585b9208d63941b3a491c202bf024779bb6e6f2adb73834c8db25e5

Observation 3551f281-1f14-42b8-9919-51b480181047 · outbound

This paper cites MMDetection: Open MMLab Detection Toolbox and Benchmark.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs MMDetection: Open MMLab Detection Toolbox and Benchmark

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:11.224870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:11.224870Z digest=sha256:043da966b13b2038879a38e6f48f5f334a2b66cdb9c62e5085df80676bcd0bcb

Observation ea12eca6-b266-45dc-be83-1cae9595b108 · outbound

This paper cites Quantifying the Carbon Emissions of Machine Learning.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Quantifying the Carbon Emissions of Machine Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:11.392668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:11.392668Z digest=sha256:d44ec25e693575dc0e2df315c3b66e7c8f35fa150ed88df0bc2122ccb5e1f574

Observation fa83d959-90c6-4d5a-91d8-f7ecd2af6b26 · outbound

This paper cites Attention Entropies.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Attention Entropies

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:23.194826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:11.624886Z digest=sha256:093c3b9797b67e460eec77a72867b186148947f27c451f3fb07909147e1688ad

Observation fb0eeb87-84ec-4df2-8f9f-f3d8bb842029 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:22.707460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:11.899321Z digest=sha256:6b3201833724b0c8790c219b4f82dd77b95c148f092d433efe8ce8fa07d10427

Observation b46e9033-6c7b-4434-83ac-9d401faf3c38 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:22.273985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:12.054837Z digest=sha256:18eaad811cd64ffb178258e803ddcc9f79bd0f1c37985c82b9d4d361cf1932bb

Observation 23b300f3-d459-43fd-81fd-970983943bcc · outbound

This paper cites Finally, we highlight the high quality of the attention maps of LUViT in the Imagenet-Segmentation dataset [Gao et al., 2022] in Section B.3.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Finally, we highlight the high quality of the attention maps of LUViT in the Imagenet-Segmentation dataset [Gao et al., 2022] in Section B.3

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:21.994766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:12.195882Z digest=sha256:b724224fb5c0f43615c8ce4ea5c1f851a429f34646d6a22c5174cf73bdcd95e0

Observation 308f5b8b-7143-4025-8fdd-d4e5232c97ad · outbound

This paper cites The results of both our reproduction of Pang et al.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs The results of both our reproduction of Pang et al

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:21.721932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:12.304770Z digest=sha256:148eadbaee9ba0d06824c968010e15e077f34c096b40a4bcde890fffed8ce19e

Observation 2c539026-45a4-408b-8d8f-6f6784ee44f9 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:21.364775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:12.514747Z digest=sha256:69b0836a9858ccf9be4dc2793b16225d343d878e2394ed21b0c8a3d068032432

Observation 5dc585a5-23fb-4816-8628-515f380fa421 · outbound

This paper cites Each reported value is an average of three training runs with three seeds, (0, 1, 2), and the subscript ± denotes the standard error for each setting.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Each reported value is an average of three training runs with three seeds, (0, 1, 2), and the subscript ± denotes the standard error for each setting

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:20.534750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:12.805484Z digest=sha256:ef7635a7557629713ab3279f2290df2121d4292e9246035b3dbfd22644cdc0a6

Observation e26f0bfe-e2bf-4b04-861c-9b75690a5919 · outbound

This paper cites A cell in the downsampled mask is assigned a value of 1 if it overlaps with the original high-resolution mask.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs A cell in the downsampled mask is assigned a value of 1 if it overlaps with the original high-resolution mask

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:20.254747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:12.974750Z digest=sha256:453b23699cd72d0b1a32e6d3e897ec0ba8f2695fcf0f8f2952c48ba1b12ff264

Observation 839f31c0-7bc6-4924-8a4c-354758aaeef4 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:19.894763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:13.145214Z digest=sha256:680b9523df8c4cbdbc7d077a1c1063a9ac0ea48518124e38ce2e1ddf55835a4c

Observation c1b92f23-9be1-49a0-89cd-15577e5ca49c · outbound

This paper cites While LUViT differs from Pang et al.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs While LUViT differs from Pang et al

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:19.505420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:13.371752Z digest=sha256:fdd2fea6e9a3edbbac641f79d465f45fc50e56fcc1d48ddcd5c4e9e6f095d453

Observation da3ffe0c-fa7e-4a2a-9b93-62a4ef5b3ec6 · outbound

This paper cites The authors quantified this alignment through demonstrating improved gradient-signal-to-noise ratio (GSNR) under the presence of the LLM block.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs The authors quantified this alignment through demonstrating improved gradient-signal-to-noise ratio (GSNR) under the presence of the LLM block

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:19.186193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:13.524735Z digest=sha256:003fff9a2c977332ecb561c34068af92126bf98f9945b91d1e98b5a5740a2703

Observation 5de44f84-dac2-4840-ac07-61247b0d0f55 · outbound

This paper cites Following up from this observation and taking inspirations from Tiwari and Shenoy [2023], Bai et al.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Following up from this observation and taking inspirations from Tiwari and Shenoy [2023], Bai et al

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:18.974743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:13.682924Z digest=sha256:47d84b23cd55e3e2a8d7f46bb229a925db2c26975a0aeb900c0595f4e9bc03ea

Observation b93c1078-e07f-4bf8-b17e-4ba73ffc0095 · outbound

This paper cites This auxiliary training objective distills the representations of the frozen-LLM-appended ViT to a vanilla ViT through a similarity loss in-between [Hinton et al., 2015].

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs This auxiliary training objective distills the representations of the frozen-LLM-appended ViT to a vanilla ViT through a similarity loss in-between [Hinton et al., 2015]

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:18.727982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:13.865294Z digest=sha256:95c443a70ff3dec06651d1f94194c3b617ec343e47026d5ab221d9064309c139

Observation d454685a-4cad-4bbe-8758-f8265300bd6a · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:17.959249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:14.464832Z digest=sha256:2e589337fd9b3f76f3fb3ba2d332c63297e4469db15ce58c3294a3fa9b42d867

Observation 9fdb8a9a-67fb-40f9-a4c6-fe9a7102f9ac · outbound

This paper cites Notably, we utilize average pooling setting instead of relying on the [CLS] token for performing classification.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Notably, we utilize average pooling setting instead of relying on the [CLS] token for performing classification

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:17.384759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:14.713999Z digest=sha256:000b15c1fb8880c3151b85436eacec15e1cd890404c8a2e186440c61a8062a56

Observation 698b20cc-0325-4398-a81c-e2335a8c7cb1 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:17.103304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:14.794778Z digest=sha256:c402650382afd2b5df4c25e10f0945cc070521cdd4143029564a9ad71c938ee2

Observation fc64965e-b720-4e62-8ec6-4f0240a71397 · outbound

This paper cites renditions.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs renditions

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:16.834523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:14.935499Z digest=sha256:5a69acf5cbb9e7c4081fc0e12f9ffaddf9067787abd8da2e5e35f201b6990174

Observation e6e964db-c461-44ad-aad6-b0501c535c97 · outbound

This paper cites Finally, the additional capacity baselines in Section 4 all have additional linear projection layers at the head, analogously with where they are placed in LUViT.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Finally, the additional capacity baselines in Section 4 all have additional linear projection layers at the head, analogously with where they are placed in LUViT

Reference 512

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:17.747453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:14.573708Z digest=sha256:78e7bc80dae51e78f283c7d231a1fde2766be3385497a1b6ad547bc787635730

Observation 7d8a0460-13f5-4f0f-ae51-22d3e18d2b07 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 768

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:18.423887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:14.072992Z digest=sha256:6c24ff09ac3b721ebb8e7d90eae0d92026da3fdb9ed55e9dc833e9a098f46485

Observation 75a915c7-eb95-4384-a911-b85a519eb630 · outbound

This paper cites Benchmarking Neural Network Robustness to Common Corruptions and Perturbations.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Benchmarking Neural Network Robustness to Common Corruptions and Perturbations

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.425779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.425779Z digest=sha256:61231b7dab629842ecff69af6921236b71dc9664bce2d102f8714b6a482b64ea

Observation 027f2cbc-e45a-4f23-a60b-184b396ade95 · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs BEiT: BERT Pre-Training of Image Transformers

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:07.315858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:07.315858Z digest=sha256:4a30b3c6d0ec992b4605dbecb889de17f81da8226a3d096ddf8c5d8ac74f201d

Observation 47bfb469-ff1f-4a19-829b-6e2b16311c60 · outbound

This paper cites Decoupled Weight Decay Regularization.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Decoupled Weight Decay Regularization

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:10.424828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:10.424828Z digest=sha256:327d7e0fbd37d90f0ae43bbc46234c60da927ad906f11230f90309ddb4c6c8b3

Observation d54a2bed-4601-40be-852d-3d6a8fecd470 · outbound

This paper cites EVEv2: Improved Baselines for Encoder-Free Vision-Language Models.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.482217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.482217Z digest=sha256:7d9e1bc51fe57dd230eceb5aeaf1ff9599ee57e41cd1ee12cdf66e28cb24a985

Observation ffa1e9e9-f158-4635-8296-6e50ca07ab6e · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.984740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.984740Z digest=sha256:9427c1582a07137ef00c774cb889b27b08dfcb85409869a9fe6eff17dc221942

Observation 5d8588d1-12b7-4214-92d8-2a3c65addc25 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:10.104743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:10.104743Z digest=sha256:89a8da9fa97bfc905a6d6f64781ece87dc7ae984ef3b8a16883a8d2dece063c8

Observation 3bb7e0bc-074e-4c44-b5fc-3728678743ab · outbound

This paper cites Noise or Signal: The Role of Image Backgrounds in Object Recognition.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Noise or Signal: The Role of Image Backgrounds in Object Recognition

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.554972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.554972Z digest=sha256:2d6782ba0d9e1ce56ca450ee1a7bbc14477f2da296afe20a2d1b7a7a66074bd9

Observation e2a04bf1-b2b9-4a7c-9aef-03e91098f537 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:06.272206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:06.272206Z digest=sha256:349608adb7a0292ca9fbd383b6d99aa3588cb293a14b1cd6966b3d856fa84336

Observation ebad85fb-f08e-4e2c-83cd-bf4aabd2f59e · outbound

This paper cites iBOT: Image BERT Pre-Training with Online Tokenizer.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs iBOT: Image BERT Pre-Training with Online Tokenizer

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:07.479266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:07.479266Z digest=sha256:24e840610fe6e9b62562034106b67b3f50a1f3f37550f12fc64267f5792e1015

Observation 8976a14b-d71a-4b71-a24b-64905332c8cf · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:06.519528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:06.519528Z digest=sha256:dd251bd6fcf6b9e5bcb929de8dd1c2b3ae8d947c9523c4c0cf5be3d8f67deebd

Observation 9c1ae321-34d2-4f7c-8dfb-f4ce852f1aee · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:07.660329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:07.660329Z digest=sha256:083e604ac2fb97bd530b54b5bc0a8f877b0e7e36707be7875a00d6aabffecc89

Observation 2f706a71-111d-4a0b-8e01-496b98b0c5c7 · outbound

This paper cites Vision as LoRA.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Vision as LoRA

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:07.858228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:07.858228Z digest=sha256:c240d9ece0a954314571cdc4d748a6dcc636358635ab7249e53e26e50644ecb5

Observation ccd954c7-0920-4a4d-97e2-fd6224cf7758 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:06.711803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:06.711803Z digest=sha256:a14749dc684e3e71fd776f99d01584211683a2e465332232bd63fdf31798f950

Observation 6d3a361d-a3ff-45e4-8a4f-53d1e34cc2da · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 4096

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:18.206814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:17:14.265854Z digest=sha256:8b7e121a847995d9437f23d1ccf9b4390d5c601ac260a434958c8e84665d6cbe

Pith citing papers

No inbound Pith citation observations are available.