Pith. sign in

Paper Citation Record · LEDGER

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs

As of 7 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2507.00754.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00754 v2

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:17:14.935499Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf57972b-3253-4d67-8cc2-8dbcf447dd24 · outbound

This paper cites Language models are few-shot learners.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Language models are few-shot learners

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:24.184422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:06.174747Z digest=sha256:60542801df245c8a24b318048ded4f6035fd160eebd3dd5f131e8305300baf7f

Observation 0d832a76-bf8a-4879-b8af-89e85431c105 · outbound

This paper cites In these works, they provided litmus tests for measuring how focused the attention patterns of particular models are and how they relate to model robustness.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs In these works, they provided litmus tests for measuring how focused the attention patterns of particular models are and how they relate to model robustness

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:22.984748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:11.783929Z digest=sha256:88fb7a32da75238e480e13971131638ce7fbc4f120c189a2de85cec8cc7c8aee

Observation 80bb878e-7eab-488d-bf54-fac98f9d577b · outbound

This paper cites the standard deviation of the accuracy values divided by the number of different seeds.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs the standard deviation of the accuracy values divided by the number of different seeds

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:20.966369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:12.633122Z digest=sha256:7344173bcd2a8dfa91f4f44bb661fc4d3f6c6d784da185aac6fd1f092cc815b9

Observation 7df17b61-bd53-4908-bfe4-0825ff4c8666 · outbound

This paper cites Frozen Transformers in Language Models Are Effective Visual Encoder Layers.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Frozen Transformers in Language Models Are Effective Visual Encoder Layers

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:17:16.472633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:06.952393Z digest=sha256:d7d73607e4a1b61bc2a5237559538851093d4c61972acfa24587a482831e0ab8

Observation 8067ec4b-4bd5-4eab-8942-b273bfb5247d · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:07.140211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:07.140211Z digest=sha256:5a9f8c35547d1a50085b88b21dd4e862debcaf8653848ca399c0a4514cf1d2da

Observation f0432910-0c1c-432f-b169-d2edd481baa3 · outbound

This paper cites Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.054862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.054862Z digest=sha256:6e97a6b4a1decbb09b48338390186c39f4d9a12817df34df72898882bd53cca0

Observation a6a28afd-2ec9-47d8-a024-7db7890eb3fc · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs DINOv2: Learning Robust Visual Features without Supervision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.132343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.132343Z digest=sha256:a08e5c83cf320afce4028842a0e884d77e1cd21c33140e7776cbd595ece809c4

Observation f5907d37-48d9-4cc7-9e57-bc8d05756ac8 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Imagenet: A large-scale hierarchical image database

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.286198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.286198Z digest=sha256:6a55ec5cb742e2830007536a26d035a2bebe3411030ff52d67d4f9a70fc46d67

Observation b3724c4e-8c25-47a4-8b1f-0025c935b7d9 · outbound

This paper cites Microsoft coco: Common objects in context.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Microsoft coco: Common objects in context

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.730823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.730823Z digest=sha256:62645cf38f429e45d4a1c6d823635e1a74c9866c45ff9404d0a6c4c05c219ee9

Observation dbb7affd-7b08-4c85-9bcd-6eb3e47e0215 · outbound

This paper cites Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language Models.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.954914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.954914Z digest=sha256:63fb529768c4e093e9f9128e4a83d5de384c83216176e69f098db9216f8b9451

Observation 54bfea88-6ae9-486d-80d8-fe4d093f67bc · outbound

This paper cites Torch.manual_seed(3407) is all you need: On the influence of random seeds in deep learning architectures for computer vision.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Torch.manual_seed(3407) is all you need: On the influence of random seeds in deep learning architectures for computer vision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.069919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.069919Z digest=sha256:cf428534ba5467ae4dcc547d16e417ff9d6440103db42c6389abeb6ccf760e6d

Observation 6a846dd8-e028-4ee5-ad8d-aa0c10fafebe · outbound

This paper cites Understanding Why Neural Networks Generalize Well Through GSNR of Parameters.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Understanding Why Neural Networks Generalize Well Through GSNR of Parameters

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.234749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.234749Z digest=sha256:4750ca502a1c9635c45ea73b66d220cc41775c690c4aca4d29709cdf9f340707

Observation a964eac8-6510-4018-b4b1-d0208ec5c971 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Distilling the Knowledge in a Neural Network

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.364748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.364748Z digest=sha256:57df05d88e5efa57f566a488665c97067b0dd2236119548d94909d06e929f864

Observation c1b79421-ce28-4c47-b15b-bc68e211d958 · outbound

This paper cites SOLO: A Single Transformer for Scalable Vision-Language Modeling.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.614748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.614748Z digest=sha256:3c52d2e75458222b598c4193493b040fc028ed3a0c8b2411db02d294d13ce910

Observation 7eb34fdb-15b0-4387-afd5-973902d3b467 · outbound

This paper cites Layer Normalization.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Layer Normalization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.764750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.764750Z digest=sha256:7e6a4f6bc641b26408795346e80d56e1ff9aaa120afe0afab12fcf6138b3061a

Observation 36fc9a21-bae4-4c9b-82ca-5261a805f6d7 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Adam: A Method for Stochastic Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:10.275076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:10.275076Z digest=sha256:0bc09a5fa787059af59abbce184b3d4d49a05825ffc4ba118ec8bd22870b1139

Observation b9ac64ec-a47d-4723-b68d-338970fc7d57 · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:10.544749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:10.544749Z digest=sha256:3126b9b017f101fa528aa6807db46e018b1d7da19396b2011cb9ae3ac6e5dbf9

Observation 330ec6c4-24c3-4029-9c70-faf41c147242 · outbound

This paper cites Deep networks with stochastic depth.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Deep networks with stochastic depth

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:23.564748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:10.764825Z digest=sha256:014f9554587d91174812fffdc2711bd16e36696fcfcca3de4423ed178995c8b4

Observation dde40791-2763-47ba-85c7-68ab4f603424 · outbound

This paper cites mixup: Beyond Empirical Risk Minimization.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs mixup: Beyond Empirical Risk Minimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:10.884728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:10.884728Z digest=sha256:98309586f35a40f2d3b13baffeb8ec513b42200a06089846d4cb298b101d1772

Observation e882166e-8e34-4822-a2a3-3639df5af93f · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:11.024571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:11.024571Z digest=sha256:3fca6d76997ea6988e2bd1fdac1135e0917a691c9ce46bef06305632b065d695

Observation 3551f281-1f14-42b8-9919-51b480181047 · outbound

This paper cites MMDetection: Open MMLab Detection Toolbox and Benchmark.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs MMDetection: Open MMLab Detection Toolbox and Benchmark

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:11.224870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:11.224870Z digest=sha256:4f1d77329fef159ecb3e6d329cce7d1229e2c0647b1614eeaf9b9921ade7f15d

Observation ea12eca6-b266-45dc-be83-1cae9595b108 · outbound

This paper cites Quantifying the Carbon Emissions of Machine Learning.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Quantifying the Carbon Emissions of Machine Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:11.392668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:11.392668Z digest=sha256:23ae949351bc540b7c39f9971335c231e1816b5961b02bd74054015db6b07cdd

Observation fa83d959-90c6-4d5a-91d8-f7ecd2af6b26 · outbound

This paper cites Attention Entropies.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Attention Entropies

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:23.194826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:11.624886Z digest=sha256:f01c5c953c283b805634c5f83b07b9f2898e522db5a41e46b8760159c63b2984

Observation fb0eeb87-84ec-4df2-8f9f-f3d8bb842029 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:22.707460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:11.899321Z digest=sha256:4f08c454eaf9dc552cf2f49c83185518b1f5d61ea2d952acf2b9662482013148

Observation b46e9033-6c7b-4434-83ac-9d401faf3c38 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:22.273985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:12.054837Z digest=sha256:7d5e9f586a329d62c64e648e18142ec8b1dfa826b8a859eb7ee5ccb5baf6ea40

Observation 23b300f3-d459-43fd-81fd-970983943bcc · outbound

This paper cites Finally, we highlight the high quality of the attention maps of LUViT in the Imagenet-Segmentation dataset [Gao et al., 2022] in Section B.3.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Finally, we highlight the high quality of the attention maps of LUViT in the Imagenet-Segmentation dataset [Gao et al., 2022] in Section B.3

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:21.994766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:12.195882Z digest=sha256:0faa5a6f417d6c350ec516f2e65300c89aede909b85b3bc27429bc8aa25c9a59

Observation 308f5b8b-7143-4025-8fdd-d4e5232c97ad · outbound

This paper cites The results of both our reproduction of Pang et al.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs The results of both our reproduction of Pang et al

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:21.721932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:12.304770Z digest=sha256:c05130bbdd886546bb55ec2539116fed177b3a2caaa4878753d988a2d8d28572

Observation 2c539026-45a4-408b-8d8f-6f6784ee44f9 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:21.364775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:12.514747Z digest=sha256:75afc96dcd5544d6cb7dc8e9bbb94c3bbad4da1ef97104d39dba8015cbffe97a

Observation 5dc585a5-23fb-4816-8628-515f380fa421 · outbound

This paper cites Each reported value is an average of three training runs with three seeds, (0, 1, 2), and the subscript ± denotes the standard error for each setting.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Each reported value is an average of three training runs with three seeds, (0, 1, 2), and the subscript ± denotes the standard error for each setting

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:20.534750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:12.805484Z digest=sha256:4d2efd5fb1a412ad82cdaa089f893127cf9749d0a1e635ace4af41c7f777a25b

Observation e26f0bfe-e2bf-4b04-861c-9b75690a5919 · outbound

This paper cites A cell in the downsampled mask is assigned a value of 1 if it overlaps with the original high-resolution mask.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs A cell in the downsampled mask is assigned a value of 1 if it overlaps with the original high-resolution mask

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:20.254747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:12.974750Z digest=sha256:f27abc9b7f8ac3977979a5aa094b580a6bfb8f5d6a467c8200c24e7017df6107

Observation 839f31c0-7bc6-4924-8a4c-354758aaeef4 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:19.894763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:13.145214Z digest=sha256:078272587ecdb46736384abe93ded8cf2d35a2a53c53ce73e2ae9d7e43cdcaa3

Observation c1b92f23-9be1-49a0-89cd-15577e5ca49c · outbound

This paper cites While LUViT differs from Pang et al.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs While LUViT differs from Pang et al

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:19.505420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:13.371752Z digest=sha256:641d256953728e6324d9952e4fc658a61915ff68c154063beddd48e4fed88d47

Observation da3ffe0c-fa7e-4a2a-9b93-62a4ef5b3ec6 · outbound

This paper cites The authors quantified this alignment through demonstrating improved gradient-signal-to-noise ratio (GSNR) under the presence of the LLM block.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs The authors quantified this alignment through demonstrating improved gradient-signal-to-noise ratio (GSNR) under the presence of the LLM block

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:19.186193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:13.524735Z digest=sha256:6f2f6fafd3dacd7e31da949412b8b18e06897888c3f604cb18c0be45e1126b74

Observation 5de44f84-dac2-4840-ac07-61247b0d0f55 · outbound

This paper cites Following up from this observation and taking inspirations from Tiwari and Shenoy [2023], Bai et al.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Following up from this observation and taking inspirations from Tiwari and Shenoy [2023], Bai et al

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:18.974743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:13.682924Z digest=sha256:031954bc02e7abeb0e1ae00989a14d39c3e10ff816e6349fa4ef1202b26ba77f

Observation b93c1078-e07f-4bf8-b17e-4ba73ffc0095 · outbound

This paper cites This auxiliary training objective distills the representations of the frozen-LLM-appended ViT to a vanilla ViT through a similarity loss in-between [Hinton et al., 2015].

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs This auxiliary training objective distills the representations of the frozen-LLM-appended ViT to a vanilla ViT through a similarity loss in-between [Hinton et al., 2015]

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:18.727982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:13.865294Z digest=sha256:8c4c32dd234d56020704eada173aed9b10370a629db78a6fcce364115d121979

Observation d454685a-4cad-4bbe-8758-f8265300bd6a · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:17.959249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:14.464832Z digest=sha256:6a082f6aeed491c72c660014dea7efefe7d8cebd9c330f0adda7320e79de1abc

Observation 9fdb8a9a-67fb-40f9-a4c6-fe9a7102f9ac · outbound

This paper cites Notably, we utilize average pooling setting instead of relying on the [CLS] token for performing classification.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Notably, we utilize average pooling setting instead of relying on the [CLS] token for performing classification

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:17.384759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:14.713999Z digest=sha256:0f25f54f96e0416f342202f090666cda5a5176e83f7d4d6de1899543fb011da9

Observation 698b20cc-0325-4398-a81c-e2335a8c7cb1 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:17.103304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:14.794778Z digest=sha256:8b06bcd0c6e7483f6aa84766b6dc417c2c37c3ebc9a94264d81b3efc2f3453d7

Observation fc64965e-b720-4e62-8ec6-4f0240a71397 · outbound

This paper cites renditions.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs renditions

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:16.834523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:14.935499Z digest=sha256:8356c79d10799b931a1c4ca327eb7e6f45c25374fa70f0eacedffc208486c7f2

Observation e6e964db-c461-44ad-aad6-b0501c535c97 · outbound

This paper cites Finally, the additional capacity baselines in Section 4 all have additional linear projection layers at the head, analogously with where they are placed in LUViT.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Finally, the additional capacity baselines in Section 4 all have additional linear projection layers at the head, analogously with where they are placed in LUViT

Reference 512

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:17.747453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:14.573708Z digest=sha256:5edb70c9b2fc3da5cca74ff0a8b83c69e8fcb1cbb4f338dd7f362f928a3e71b4

Observation 7d8a0460-13f5-4f0f-ae51-22d3e18d2b07 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 768

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:18.423887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:14.072992Z digest=sha256:8f712d1c2f1ac8439eceb86b2c26d01e1ad0a5db7fb372bb58813e9257c9d043

Observation 75a915c7-eb95-4384-a911-b85a519eb630 · outbound

This paper cites Benchmarking Neural Network Robustness to Common Corruptions and Perturbations.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Benchmarking Neural Network Robustness to Common Corruptions and Perturbations

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.425779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.425779Z digest=sha256:b9a7019cdc3dc99a2968efec9c876a3d13803b0dae4a6c4644434d76edd78cd7

Observation 027f2cbc-e45a-4f23-a60b-184b396ade95 · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs BEiT: BERT Pre-Training of Image Transformers

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:07.315858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:07.315858Z digest=sha256:18bde7f9d39e058d670572065c4db65e4a5c4c577905d9fac117e04ba553587a

Observation 47bfb469-ff1f-4a19-829b-6e2b16311c60 · outbound

This paper cites Decoupled Weight Decay Regularization.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Decoupled Weight Decay Regularization

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:10.424828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:10.424828Z digest=sha256:5733cb82fc88c46abfe1fec5407efb75e02f452eb5ccb0fa34375302c0a6a669

Observation d54a2bed-4601-40be-852d-3d6a8fecd470 · outbound

This paper cites EVEv2: Improved Baselines for Encoder-Free Vision-Language Models.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.482217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.482217Z digest=sha256:e9e401ef9d0248d9fd2f470f749c264533e41b5f044a3041982eb5368b314352

Observation ffa1e9e9-f158-4635-8296-6e50ca07ab6e · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.984740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.984740Z digest=sha256:6acc0388d124587834bc74e639345b968ca0293a5b06c525694e80940e122feb

Observation 5d8588d1-12b7-4214-92d8-2a3c65addc25 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:10.104743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:10.104743Z digest=sha256:15a780f107c7f6e6ae93010f42996b75e398884c7b80090e5f0c8fe330aba81d

Observation 3bb7e0bc-074e-4c44-b5fc-3728678743ab · outbound

This paper cites Noise or Signal: The Role of Image Backgrounds in Object Recognition.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Noise or Signal: The Role of Image Backgrounds in Object Recognition

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.554972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.554972Z digest=sha256:473e9bbf0157c9f101e74af0cebebe9e421ff6f193cca77d009c9643559266ad

Observation e2a04bf1-b2b9-4a7c-9aef-03e91098f537 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:06.272206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:06.272206Z digest=sha256:98f7caaf9eb7f835c67d45631a3b6a91e726a09ff629dd8e29fcd16543a6e666

Observation ebad85fb-f08e-4e2c-83cd-bf4aabd2f59e · outbound

This paper cites iBOT: Image BERT Pre-Training with Online Tokenizer.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs iBOT: Image BERT Pre-Training with Online Tokenizer

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:07.479266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:07.479266Z digest=sha256:2205c2cf379ac89774aa4e17ef80ceed2a1fa6caec9fbc6c8651d53c133c289f

Observation 8976a14b-d71a-4b71-a24b-64905332c8cf · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:06.519528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:06.519528Z digest=sha256:e88732f0f6dc9462f4966eb723f5c8024632d3192e75884caa338e05dca1f868

Observation 9c1ae321-34d2-4f7c-8dfb-f4ce852f1aee · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:07.660329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:07.660329Z digest=sha256:f2b1f9a59426e351d40bb1cfcc5926840d2bc9a5f2c9b2e57116146c6bba30ef

Observation 2f706a71-111d-4a0b-8e01-496b98b0c5c7 · outbound

This paper cites Vision as LoRA.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Vision as LoRA

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:07.858228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:07.858228Z digest=sha256:9544034fd15310068838b0dc8f2620b077a21b0504c7f749f4504703565befc9

Observation ccd954c7-0920-4a4d-97e2-fd6224cf7758 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:06.711803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:06.711803Z digest=sha256:ef940d6b8a19217c814bdc99e3396586d3421f7624b433a81b58c0e367b60698

Observation 6d3a361d-a3ff-45e4-8a4f-53d1e34cc2da · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 4096

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:18.206814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:17:14.265854Z digest=sha256:c47462a88cf57a437e6a5ef87d14d4fe77e4871ca9d4bb5cf8f02ca840d7bb3a

Pith citing papers

No inbound Pith citation observations are available.