Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:17:14.935499Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2507.00754.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:17:14.935499Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cf57972b-3253-4d67-8cc2-8dbcf447dd24 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Language models are few-shot learners
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0d832a76-bf8a-4879-b8af-89e85431c105 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs In these works, they provided litmus tests for measuring how focused the attention patterns of particular models are and how they relate to model robustness
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80bb878e-7eab-488d-bf54-fac98f9d577b · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs the standard deviation of the accuracy values divided by the number of different seeds
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7df17b61-bd53-4908-bfe4-0825ff4c8666 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Frozen Transformers in Language Models Are Effective Visual Encoder Layers
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8067ec4b-4bd5-4eab-8942-b273bfb5247d · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0432910-0c1c-432f-b169-d2edd481baa3 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6a28afd-2ec9-47d8-a024-7db7890eb3fc · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs DINOv2: Learning Robust Visual Features without Supervision
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5907d37-48d9-4cc7-9e57-bc8d05756ac8 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Imagenet: A large-scale hierarchical image database
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3724c4e-8c25-47a4-8b1f-0025c935b7d9 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Microsoft coco: Common objects in context
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbb7affd-7b08-4c85-9bcd-6eb3e47e0215 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54bfea88-6ae9-486d-80d8-fe4d093f67bc · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Torch.manual_seed(3407) is all you need: On the influence of random seeds in deep learning architectures for computer vision
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a846dd8-e028-4ee5-ad8d-aa0c10fafebe · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Understanding Why Neural Networks Generalize Well Through GSNR of Parameters
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a964eac8-6510-4018-b4b1-d0208ec5c971 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Distilling the Knowledge in a Neural Network
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1b79421-ce28-4c47-b15b-bc68e211d958 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs SOLO: A Single Transformer for Scalable Vision-Language Modeling
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7eb34fdb-15b0-4387-afd5-973902d3b467 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Layer Normalization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36fc9a21-bae4-4c9b-82ca-5261a805f6d7 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Adam: A Method for Stochastic Optimization
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9ac64ec-a47d-4723-b68d-338970fc7d57 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 330ec6c4-24c3-4029-9c70-faf41c147242 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Deep networks with stochastic depth
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dde40791-2763-47ba-85c7-68ab4f603424 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs mixup: Beyond Empirical Risk Minimization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e882166e-8e34-4822-a2a3-3639df5af93f · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs PyTorch: An Imperative Style, High-Performance Deep Learning Library
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3551f281-1f14-42b8-9919-51b480181047 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs MMDetection: Open MMLab Detection Toolbox and Benchmark
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea12eca6-b266-45dc-be83-1cae9595b108 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Quantifying the Carbon Emissions of Machine Learning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa83d959-90c6-4d5a-91d8-f7ecd2af6b26 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Attention Entropies
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fb0eeb87-84ec-4df2-8f9f-f3d8bb842029 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b46e9033-6c7b-4434-83ac-9d401faf3c38 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 23b300f3-d459-43fd-81fd-970983943bcc · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Finally, we highlight the high quality of the attention maps of LUViT in the Imagenet-Segmentation dataset [Gao et al., 2022] in Section B.3
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 308f5b8b-7143-4025-8fdd-d4e5232c97ad · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs The results of both our reproduction of Pang et al
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c539026-45a4-408b-8d8f-6f6784ee44f9 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5dc585a5-23fb-4816-8628-515f380fa421 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Each reported value is an average of three training runs with three seeds, (0, 1, 2), and the subscript ± denotes the standard error for each setting
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e26f0bfe-e2bf-4b04-861c-9b75690a5919 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs A cell in the downsampled mask is assigned a value of 1 if it overlaps with the original high-resolution mask
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 839f31c0-7bc6-4924-8a4c-354758aaeef4 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c1b92f23-9be1-49a0-89cd-15577e5ca49c · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs While LUViT differs from Pang et al
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da3ffe0c-fa7e-4a2a-9b93-62a4ef5b3ec6 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs The authors quantified this alignment through demonstrating improved gradient-signal-to-noise ratio (GSNR) under the presence of the LLM block
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5de44f84-dac2-4840-ac07-61247b0d0f55 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Following up from this observation and taking inspirations from Tiwari and Shenoy [2023], Bai et al
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b93c1078-e07f-4bf8-b17e-4ba73ffc0095 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs This auxiliary training objective distills the representations of the frozen-LLM-appended ViT to a vanilla ViT through a similarity loss in-between [Hinton et al., 2015]
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d454685a-4cad-4bbe-8758-f8265300bd6a · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9fdb8a9a-67fb-40f9-a4c6-fe9a7102f9ac · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Notably, we utilize average pooling setting instead of relying on the [CLS] token for performing classification
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 698b20cc-0325-4398-a81c-e2335a8c7cb1 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc64965e-b720-4e62-8ec6-4f0240a71397 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs renditions
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e6e964db-c461-44ad-aad6-b0501c535c97 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Finally, the additional capacity baselines in Section 4 all have additional linear projection layers at the head, analogously with where they are placed in LUViT
Reference 512
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d8a0460-13f5-4f0f-ae51-22d3e18d2b07 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work
Reference 768
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 75a915c7-eb95-4384-a911-b85a519eb630 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Benchmarking Neural Network Robustness to Common Corruptions and Perturbations
Reference 2009
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 027f2cbc-e45a-4f23-a60b-184b396ade95 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs BEiT: BERT Pre-Training of Image Transformers
Reference 2010
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47bfb469-ff1f-4a19-829b-6e2b16311c60 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Decoupled Weight Decay Regularization
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d54a2bed-4601-40be-852d-3d6a8fecd470 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffa1e9e9-f158-4635-8296-6e50ca07ab6e · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d8588d1-12b7-4214-92d8-2a3c65addc25 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs SGDR: Stochastic Gradient Descent with Warm Restarts
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bb7e0bc-074e-4c44-b5fc-3728678743ab · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Noise or Signal: The Role of Image Backgrounds in Object Recognition
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2a04bf1-b2b9-4a7c-9aef-03e91098f537 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs LLaMA: Open and Efficient Foundation Language Models
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebad85fb-f08e-4e2c-83cd-bf4aabd2f59e · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs iBOT: Image BERT Pre-Training with Online Tokenizer
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8976a14b-d71a-4b71-a24b-64905332c8cf · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c1ae321-34d2-4f7c-8dfb-f4ce852f1aee · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f706a71-111d-4a0b-8e01-496b98b0c5c7 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Vision as LoRA
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccd954c7-0920-4a4d-97e2-fd6224cf7758 · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d3a361d-a3ff-45e4-8a4f-53d1e34cc2da · outbound
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work
Reference 4096
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.