Pith. sign in

Paper Citation Record · LEDGER

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training

As of 18 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2411.11927.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11927 v3

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:38:19.066615Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:20:50.578691Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy39
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5fc2996f-d8e4-4fdc-8e1c-b26acbc91c16 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.787439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.787439Z digest=sha256:fa74688783f14805fdb88faafc8d7ee849f5a8990e0f52b9311f2eacd7f4644b

Observation 771fe683-689f-441d-b6e9-682a5f800c57 · outbound

This paper cites GPT-4 Technical Report.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.792629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.792629Z digest=sha256:919e4e11604e409950cd590664dd835e172cc186262756240ae70add947e12d6

Observation 660777d5-cb9c-4b98-b151-3a197f179d4e · outbound

This paper cites Qwen Technical Report.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.797615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.797615Z digest=sha256:af273d5607b5ce7ecc5e110dca1ba1e446592f60fd8ea46ee77a14290ee8a2c5

Observation ffdf892f-bd6e-4a73-abd7-c55e88b008fd · outbound

This paper cites Is a 3D-Tokenized LLM the Key to Reliable Autonomous Driving?.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Is a 3D-Tokenized LLM the Key to Reliable Autonomous Driving?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.802194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.802194Z digest=sha256:47822592f867be0f218aff1db7a0673546a05bf1af2c85391e1502a9ec1c471e

Observation 4c869d3f-597c-4e61-92cc-3786a74a9fb0 · outbound

This paper cites Ar- trackv2: Prompting autoregressive tracker where to look and how to describe.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Ar- trackv2: Prompting autoregressive tracker where to look and how to describe

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.991922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.806733Z digest=sha256:1be1775d7110864ed7a4dc7bebc03da064d3bb7edac61b0844d44e7803434376

Observation c4c56f93-3de8-484b-a1ad-86a340e98125 · outbound

This paper cites LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.810790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.810790Z digest=sha256:2554122a947cb2f30c5e48a209f6dfd5c2223a86fc3b82820f70fb63b0b0953e

Observation 5e25dad2-cdbc-4c67-a9af-08fa36885ddf · outbound

This paper cites Language models are few-shot learners.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Language models are few-shot learners

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.981028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.815209Z digest=sha256:2287a1ae66f3981eb464ff2dc0594b7aa3f7bdbb9e47af9a874ea4396cc66c1f

Observation 0c0df006-4258-4b48-ba93-2aa5bcbf47b9 · outbound

This paper cites Cross-lingual and multilingual clip.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Cross-lingual and multilingual clip

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.968846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.819267Z digest=sha256:33541ba3c9f9359926bf6f77f06d68f970c9bbc84bb603534faff1bdf015c01c

Observation f2176700-0419-4c48-b9f3-f506a023749c · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.823234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.823234Z digest=sha256:74ca46f36d426a427647e4ee12f05ead89ce8afe8532f31f712e45e567796ee2

Observation 99a2709e-9a1a-4d99-890b-3cf78d420f12 · outbound

This paper cites Pali: A jointly- scaled multilingual language-image model.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Pali: A jointly- scaled multilingual language-image model

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.954196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.827882Z digest=sha256:28c0f2669670be1b3444882aedea8123fa9891f842a4a2e4107bbfa9c55cd47c

Observation d1b03554-275a-471a-bd59-81be51e97053 · outbound

This paper cites Altclip: Altering the language encoder in clip for extended language capabilities.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Altclip: Altering the language encoder in clip for extended language capabilities

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.941293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.833993Z digest=sha256:f732f86b8fbcf7ebab3848cb9d00bb8e6a1fd15c46e4d97b8931e498d8028e40

Observation 818f101b-df01-47cb-9ef8-8d580dbea403 · outbound

This paper cites Maskclip: Masked self-distillation advances contrastive language-image pretraining.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Maskclip: Masked self-distillation advances contrastive language-image pretraining

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.927385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.840322Z digest=sha256:31dc168ed9972c216fd7da47c17edde37bbc1b55f7fccd8072298753488372fd

Observation d7465e84-03ff-4c58-ba74-e7446b4560ab · outbound

This paper cites An image is worth 16x16 words: Transform- ers for image recognition at scale.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training An image is worth 16x16 words: Transform- ers for image recognition at scale

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.906995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.845096Z digest=sha256:8b71a338f907d0341a928069f2736f1c2eb51359bf81b17a4a400533da2a1f7d

Observation 10bec2c1-eddb-469a-9391-df7cae57fefa · outbound

This paper cites Improving clip training with language rewrites.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Improving clip training with language rewrites

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.891952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.849417Z digest=sha256:080e3c30bb5c339bcabae58166f65f81a71202703804b3303323c51b2681446d

Observation 2dc147cf-df78-40ce-8f07-695493844f6b · outbound

This paper cites Pyramidclip: Hierarchical feature alignment for vision-language model pretraining.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Pyramidclip: Hierarchical feature alignment for vision-language model pretraining

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.853452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.853452Z digest=sha256:a9c4646306be278e00b9ee707062371154e5c085c46e46078b0d221c0faaad31

Observation f562151e-0523-4909-994a-422254de4be5 · outbound

This paper cites Softclip: Softer cross-modal alignment makes clip stronger.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Softclip: Softer cross-modal alignment makes clip stronger

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.873648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.857225Z digest=sha256:5c543be412ed4a6fe6efaa35626c4b27db3cf58495bdc5fca9ca5f873a74cbc2

Observation a44e35a2-0ca5-42d9-9767-00719897ee77 · outbound

This paper cites Hiclip: Contrastive language-image pre- training with hierarchy-aware attention.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Hiclip: Contrastive language-image pre- training with hierarchy-aware attention

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.861679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.861394Z digest=sha256:82feaaa2e19efbeaf9e2e9e17194e355bd513d5117b7d4b57e82c706647af866

Observation f8b962ac-6dc7-4488-a66d-56777c37f4a0 · outbound

This paper cites Sugarcrepe: Fixing hack- able benchmarks for vision-language compositionality.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Sugarcrepe: Fixing hack- able benchmarks for vision-language compositionality

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.849945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.865235Z digest=sha256:3ef7c4b245d387e93baec09004575efcf02b602c619a3bc29c002dbea8eeb226

Observation 8a850156-b50c-410f-a52f-a7e55f11e682 · outbound

This paper cites Openclip, 2021.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Openclip, 2021

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.837127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.869235Z digest=sha256:6c48ff8c4348859876dfd797aa7bca4ec1678d3a0ed6a660064ca90eccd92d7d

Observation cd64749d-6852-4338-bceb-c524d5dfda4c · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Scaling up visual and vision-language representation learning with noisy text supervision

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.824649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.873432Z digest=sha256:080397647f766461634e7941b6e2b08d439b9094721bf020482f63f32ebd2522

Observation 00cb0a3b-3899-409a-b0e7-2a1a893c0f8f · outbound

This paper cites Mistral 7B.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Mistral 7B

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.878448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.878448Z digest=sha256:50b233efaaef9f7bd701b1fdbb264d9d3739f429cc3bf9723ded0faeb3bd0f62

Observation 8d033562-e4e6-40c2-97d8-eb4b9b788d4f · outbound

This paper cites Scaling Sentence Embeddings with Large Language Models.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Scaling Sentence Embeddings with Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.882734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.882734Z digest=sha256:47f1e138f442eaf7ada83c4e8258297eb2adf8a2161a61039e581a7b729083d3

Observation 9523231a-5fe3-433f-b323-a0f5e73eb07c · outbound

This paper cites Misalign, contrast then distill: Rethinking misalign- ments in language-image pre-training.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Misalign, contrast then distill: Rethinking misalign- ments in language-image pre-training

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.812134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.886495Z digest=sha256:e36cbe0875eb5dba0ade88b7152f722644d9ce25de17456ae7c5c77c6d420046

Observation 03a6eaf3-1f33-47d8-b5fd-cbadd5602711 · outbound

This paper cites Jina CLIP: Your CLIP Model Is Also Your Text Retriever.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Jina CLIP: Your CLIP Model Is Also Your Text Retriever

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.890177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.890177Z digest=sha256:7776be401ac40586d2b9903de4eb32244611c0fff059d05061cc8c0bd4f7bd2c

Observation df8046fa-aa08-4e92-9781-8bcb8859ac51 · outbound

This paper cites VeCLIP: Improving CLIP Training via Visual-enriched Captions.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training VeCLIP: Improving CLIP Training via Visual-enriched Captions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.894281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.894281Z digest=sha256:ba0b224cdf01bf911e599a1ad5744b99caea832bf795b06cf8dae2f602cd9a9f

Observation e468e0e3-9f0b-4f11-9434-debbc4ba3d07 · outbound

This paper cites Uni- clip: Unified framework for contrastive language-image pre- training.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Uni- clip: Unified framework for contrastive language-image pre- training

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.800689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.898476Z digest=sha256:5d9c41362ad5e559635bbd6dc9d695033b5d832fc9ebfe7db2ae69d28b7df843

Observation e9c16874-4ff8-4c72-9d0b-036dbe818bdc · outbound

This paper cites Gecko: Versatile Text Embeddings Distilled from Large Language Models.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Gecko: Versatile Text Embeddings Distilled from Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.902333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.902333Z digest=sha256:cb1d5fbc80076fee9157f65ae493c65c7a6c330c159754fe52c37c72acab3d88

Observation 10ee5516-cff1-45b8-be52-193017adfe4b · outbound

This paper cites Meta-task prompting elicits embedding from large language models.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Meta-task prompting elicits embedding from large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.788710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.906723Z digest=sha256:39e00ff4f57482fee12e774e7c753fa2a3760c1c24794a4cedc8b799da903e37

Observation 6eccc694-931a-4ab0-8990-fbe6fecc8037 · outbound

This paper cites Scene graph generation: A comprehensive survey.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Scene graph generation: A comprehensive survey

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.775909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.910853Z digest=sha256:d98527af10429cfb90c92b0a38a2e078325f60758f4bba4fcfef7cb259ed24e2

Observation 16d083a4-7b19-4f7e-861f-e3aa071eb484 · outbound

This paper cites Grounded language- image pre-training.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Grounded language- image pre-training

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.764217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.914822Z digest=sha256:eb826b6e77c4d444ded919b6e3b68eacc6037e4971680ed2a580c9a243557b84

Observation c291c074-5c93-4778-b7b3-73303cae886f · outbound

This paper cites An inverse scaling law for clip training.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training An inverse scaling law for clip training

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.752353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.919615Z digest=sha256:db0bff60789f78fb05d1a96651969bcdd148c58448627fd60f25454493ec4f33

Observation ad7acedf-ad08-4128-aa9b-576f699a806f · outbound

This paper cites Scene graph generation from objects, phrases and region captions.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Scene graph generation from objects, phrases and region captions

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.741521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.923463Z digest=sha256:9bf95bb9d204a98618b40f174b0a0bc81b04f1a4ecbf45bf65e712dece618766

Observation a00fbd1f-04a6-4e30-bb3a-28cd27b1015b · outbound

This paper cites Supervi- sion exists everywhere: A data efficient contrastive language- image pre-training paradigm.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Supervi- sion exists everywhere: A data efficient contrastive language- image pre-training paradigm

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.729879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.927328Z digest=sha256:c89d058dd7717a0b2b72df76b209ff37deb39ff2dc79b79264ce42ee69afc9f4

Observation 27244eb8-92a5-4a14-9c51-e2409d99e880 · outbound

This paper cites Scaling language-image pre-training via masking.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Scaling language-image pre-training via masking

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.931234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.931234Z digest=sha256:efd2bc9f087e3014f99a4861d0b93db64b3af8898cc19a911aa9957097351b39

Observation 7358f3c8-a38c-4ecc-ac23-1f39083baa42 · outbound

This paper cites Language quantized autoencoders: Towards unsupervised text-image alignment.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Language quantized autoencoders: Towards unsupervised text-image alignment

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.709761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.935183Z digest=sha256:89b9e4d2034bad373efd911c375db6185122cd1eea73fcf83534e2a84d540e4c

Observation 07d857dd-9ef6-473b-aeff-8efc72d12d18 · outbound

This paper cites MLLMs-Augmented Visual-Language Representation Learning.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training MLLMs-Augmented Visual-Language Representation Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.940013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.940013Z digest=sha256:688d5e8683db8bdcdf13011725ec28e2d8a760ddbbdc35aa34070c2940a492f1

Observation eb291ae4-5fcb-422f-bd0d-46559506bd91 · outbound

This paper cites Decoupled Weight Decay Regularization.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Decoupled Weight Decay Regularization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.944402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.944402Z digest=sha256:c314c266b66a42fcde1590b23348b4d0bd76ff42f8ec4379499869c6c0f5bf8f

Observation 0686173c-cdf0-4d4b-b1df-3615ec2ce85e · outbound

This paper cites Slip: Self-supervision meets language-image pre- training.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Slip: Self-supervision meets language-image pre- training

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.696620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.948730Z digest=sha256:c48db9a76b99192d79176cabfdc667245db5a27e432bf49f5aa50a104399cfee

Observation ecf32336-bbed-4254-bbdc-e17cb97b2bcf · outbound

This paper cites Generative Representational Instruction Tuning.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Generative Representational Instruction Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.952619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.952619Z digest=sha256:8b5359e5224c07a3c22c85a1f563fc79cc28f50aa16739394e45b3ef4df80047

Observation e9fcf66c-1c73-49af-a6c7-910c486ad503 · outbound

This paper cites Docci: De- scriptions of connected and contrasting images.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Docci: De- scriptions of connected and contrasting images

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.683887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.956758Z digest=sha256:4fe64591ab9dac663547305bc9c23ea3cf7b0cef7a3b26b91164b28265fd594c

Observation 37cbdc5f-e82c-4926-8245-1520109e40ca · outbound

This paper cites Learning transferable visual models from natural language supervision.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Learning transferable visual models from natural language supervision

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.670928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.960633Z digest=sha256:d4136b604b2ce8f1a2b95e0032e438c75a51178a87b6bcf4f4543a422cf91fad

Observation 5d1b15dd-82f4-4705-b955-ae8943f16fcc · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training High-resolution image synthesis with latent diffusion models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.964348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.964348Z digest=sha256:6f421c6c72e7b4eb22f17461d18b5d3788a440bc475e11d38e7b5291b396f3ba

Observation 00ea4c14-7c8a-47b0-b0eb-3be3273f8e55 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.648904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.968709Z digest=sha256:27f4dca32d37b885df721ac560a6e2be8ebb52182b8c7eab4d4298b89155508e

Observation b94bf483-407b-4f26-b040-3338f133aeb3 · outbound

This paper cites Repetition Improves Language Model Embeddings.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Repetition Improves Language Model Embeddings

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.973434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.973434Z digest=sha256:c416444763fcc1c90b89d6b0e30781c84bc3924662d02f8357c936a901549c39

Observation 1b1a5244-7d1c-4258-b856-f9a6ca575b3c · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.979973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.979973Z digest=sha256:28fe140cde07218e32cdc6ac55d58b5d3008fc9c009a2c1c4747b404a9d48d76

Observation 64e4cd87-453e-4e07-aac4-042fc769be0f · outbound

This paper cites Crossmodal-3600: A massively multilingual multi- modal evaluation dataset.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Crossmodal-3600: A massively multilingual multi- modal evaluation dataset

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.635017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.985183Z digest=sha256:2d9741230c44508be5691af9d299489036cdba645066d253afca219255af8e58

Observation a603a36f-318c-49e6-9fb0-e34cf50d15a4 · outbound

This paper cites Winoground: Probing vision and language models for visio- linguistic compositionality.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Winoground: Probing vision and language models for visio- linguistic compositionality

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.620165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.989037Z digest=sha256:407bbfe3f0f54c43fe74439c5f932cf2c4efa355bbca1db197d43cfb0cf886f3

Observation 49222733-a491-46c9-accc-61f176f490a4 · outbound

This paper cites Stablerep: Synthetic images from text-to- image models make strong visual representation learners.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Stablerep: Synthetic images from text-to- image models make strong visual representation learners

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.608271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:18.992900Z digest=sha256:2585def4524200f30e8ee8cf55d9dd377427df878b585aee4792efbbb8fa715c

Observation 3d212d76-ad4e-4472-a9a6-86c71f9e7b44 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:18.997000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:18.997000Z digest=sha256:d3458c6838d69e6423d3723f93a4498eb9556093a755ba87c665b674b8a91a43

Observation 181ea0ff-5cca-4aec-aaee-a07ffa20816e · outbound

This paper cites Multimodal few-shot learning with frozen language models.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Multimodal few-shot learning with frozen language models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.596000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:19.001300Z digest=sha256:b8e8600cd99c6c8d504112a44e8c09b21756071fe2420efe628a2467dd11e28b

Observation ee888382-854d-4e60-8872-239e223ef6e8 · outbound

This paper cites NLLB-CLIP -- train performant multilingual image retrieval model on a budget.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training NLLB-CLIP -- train performant multilingual image retrieval model on a budget

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:19.006394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:19.006394Z digest=sha256:7bbff95d149939fac7e318a5a783737e1314e6f188a471cac6e1b0f932dbed59

Observation b07fbd1e-0b69-4a54-bc5e-52af05dbb357 · outbound

This paper cites Improving Text Embeddings with Large Language Models.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Improving Text Embeddings with Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:19.011691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:19.011691Z digest=sha256:e4191255704f93b299fa4fe5159f6f895ba23c1d21f86db1a2c6a6a4a6333eb8

Observation ef407a48-25e3-4da7-8ac7-cbc02f5c9153 · outbound

This paper cites Groupvit: Semantic segmentation emerges from text supervision.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Groupvit: Semantic segmentation emerges from text supervision

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.584093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:19.017503Z digest=sha256:c177c9a2d6800370a8dbebcee03fcd2ff402ae93794336594cd70f478971f89a

Observation 5ba97f96-d30f-4d48-8a22-a1c68accd075 · outbound

This paper cites Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:19.021241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:19.021241Z digest=sha256:ba17bb3e1ac4107fc4ca9d59fae8823dbbb9c694aa1324fd53bd87b9df1d6067

Observation 0d6bbb8b-f28e-4df6-b230-c09b5f4bffa4 · outbound

This paper cites Alip: Adaptive language-image pre-training with synthetic caption.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Alip: Adaptive language-image pre-training with synthetic caption

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.572338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:19.025421Z digest=sha256:9517ac4d8db319c4f4cc80645c5d8a9b84577d328b12800a80f405eee18fef92

Observation 08a0c0fb-f71e-41cd-9394-4ac6a690eb78 · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:19.029722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:19.029722Z digest=sha256:8cbe26611b32cf7eb5e5373978907bb09d049951caf38451666329f43a26007a

Observation b9688f0f-3c42-4e1c-8b41-f20410f578c1 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:19.034360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:19.034360Z digest=sha256:576de026db7d103c93bce03cabf701f61b2b07d6f7b5a3def2cae973ce9cbf58

Observation 6700abaa-d14e-4096-b8d4-80892ca4f79f · outbound

This paper cites Spae: Semantic pyramid autoencoder for multimodal generation with frozen llms.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Spae: Semantic pyramid autoencoder for multimodal generation with frozen llms

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.456970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:19.038843Z digest=sha256:2901abf4575b78e93b52cc2a972607d150dd986ba37f2303e2465da58691addd

Observation 46fbb959-e44a-465c-a66f-4232af836c81 · outbound

This paper cites Lit: Zero-shot transfer with locked-image text tuning.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Lit: Zero-shot transfer with locked-image text tuning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.445829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:19.042838Z digest=sha256:41bc4c2012049654ebe9a237fde44bad1f79ecd1010e3e1cc7e7b3cfb28bd069

Observation 24684c16-7b27-4655-9deb-003400c7547c · outbound

This paper cites Sigmoid loss for language image pre-training.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Sigmoid loss for language image pre-training

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.433717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:19.046846Z digest=sha256:d4f249cd6ae873ce293b999f1df4ed3797f2bcd3ae160d463f5a29a10bef083e

Observation 0a8f4552-a5fc-4809-bdf8-9683dbfab0d1 · outbound

This paper cites Simple techniques for enhancing sentence embeddings in generative language models.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Simple techniques for enhancing sentence embeddings in generative language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.420132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:19.051036Z digest=sha256:f08ee94df6d50d5882443225ec648bc3b4c99eda5f74aa41ed5d3ee682caa355

Observation 688377fb-d2d8-4516-b32b-70dc64e34dc9 · outbound

This paper cites Long-clip: Unlocking the long-text capability of clip.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Long-clip: Unlocking the long-text capability of clip

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.406652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:19.054747Z digest=sha256:8d684f3f092453fc00d9290911a7680f7730410888340f2d8035f174f40bc99f

Observation d5678b42-79b4-4865-aee6-2f0a69950a5f · outbound

This paper cites Dreamlip: Language- image pre-training with long captions.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Dreamlip: Language- image pre-training with long captions

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.394875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:19.058927Z digest=sha256:9383796fa925ecfde70546e60a5e6cb0d353c9508079532258b0403812daf3b5

Observation 31db733f-781e-4cc3-a199-1dcc8e19c9c8 · outbound

This paper cites Beyond text: Frozen large language models in visual signal comprehension.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Beyond text: Frozen large language models in visual signal comprehension

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:38:19.382821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:19.062778Z digest=sha256:edbef1b6dff418e0d092275c72675e8ec8fa0d53aa8d8620e1bb7fe3bc235160

Observation bbb638db-a5df-4c67-a476-1f77e401f0d7 · outbound

This paper cites yi”. After thinking step by step, the category of the main object in this image means in just one word:.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training yi”. After thinking step by step, the category of the main object in this image means in just one word:

Reference 65

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T18:38:19.370466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T18:38:19.066615Z digest=sha256:fb9cd3b8fd55297908cab7d6ec7109e28ea5e266a6a5da279733510b21bd9830

Pith citing papers

Observation dccbc202-eb71-431c-a90a-e929ada7b54e · inbound

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs cites this paper.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:50.578691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:50.578691Z digest=sha256:047a8e971701377986c2fc5284fb19232150b0636d2a593206ae0c66dc525986