Pith. sign in

Paper Citation Record · LEDGER

Visual Lexicon: Rich Image Features in Language Space

As of 12 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 1 inbound Pith citation observation for arXiv:2412.06774.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.06774 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:22:55.635316Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:54:09.796839Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T18:54:12.830432Z

Reference resolution

89 of 89 outbound references displayed

  • verified exact0
  • verified fuzzy47
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa53fbcf-cfba-451d-b919-ac685e9b21d2 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Visual Lexicon: Rich Image Features in Language Space Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.902029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.902029Z digest=sha256:6cab5fcd87d8fccb42d94593a312106541b1e2a0a2f95a5bb555800c23f55f3e

Observation df5898d1-f414-46de-844d-3616b039a90c · outbound

This paper cites BEit: BERT pre-training of image transformers.

Visual Lexicon: Rich Image Features in Language Space BEit: BERT pre-training of image transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.911285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.911285Z digest=sha256:e021698d8be1fe7932b0ba54da674b3564aa5ad387f92715b579b429c354c6a6

Observation 8110e71e-2571-433c-812b-960dd4e6a763 · outbound

This paper cites Label-efficient se- mantic segmentation with diffusion models.

Visual Lexicon: Rich Image Features in Language Space Label-efficient se- mantic segmentation with diffusion models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.918280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.918280Z digest=sha256:bf312b0141b91ff4716f792779f36bf7260f4d5a037192d7584e2c714abd6f0d

Observation 11a12ab5-d0aa-4f21-b0e2-42b25e537971 · outbound

This paper cites Generalized denoising auto-encoders as generative models.

Visual Lexicon: Rich Image Features in Language Space Generalized denoising auto-encoders as generative models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.925181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.925181Z digest=sha256:e1390ec91ae8a73db6abc58f605411fbf6d6a57c3b9f99da52c9203290592154

Observation 2bf1a14f-f047-4925-b72f-44ddb677ce87 · outbound

This paper cites Improving image generation with better captions.

Visual Lexicon: Rich Image Features in Language Space Improving image generation with better captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.931763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.931763Z digest=sha256:9c2fa9e755cb54ef852b0de1b2a16b48e821046072dd62a02b65d788f34eb715

Observation 187cf068-12cc-4afb-89fd-f1b572370631 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Visual Lexicon: Rich Image Features in Language Space PaliGemma: A versatile 3B VLM for transfer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.939829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.939829Z digest=sha256:3218bb37a02d2d43566b39ab74be32d57f08900ebfeeabfcbe6a59cddd232d65

Observation a7233933-ab07-44fb-91cf-bdef1edd0a01 · outbound

This paper cites Language Models are Few-Shot Learners.

Visual Lexicon: Rich Image Features in Language Space Language Models are Few-Shot Learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.949489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.949489Z digest=sha256:619b03322f96a51705347cd349c9ae0f36e31ab8583de06561173e2d0bd3222e

Observation 34f3191c-b220-4d09-b037-40fae470963c · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Visual Lexicon: Rich Image Features in Language Space Emerg- ing properties in self-supervised vision transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.957417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.957417Z digest=sha256:ad66e57999e28a1403f0a58fe56798ce4e069a60e21c88fd9dff56dabd1c1f56

Observation 1cf3a614-1b2c-4dd0-b362-f39cbf9c9fa8 · outbound

This paper cites Generative pre- training from pixels.

Visual Lexicon: Rich Image Features in Language Space Generative pre- training from pixels

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.966657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.966657Z digest=sha256:d0c1a9c1ce7439874a21fee3fc612daddf5151c11bc4c1e1f34d39985718b019

Observation ed4090ed-fc6a-48d9-acaf-358c37ae14e3 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Visual Lexicon: Rich Image Features in Language Space Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.975741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.975741Z digest=sha256:3a352852e17fd343713d603f4808d05abed1a936e505ea2dd086e23acac0b678

Observation 4471d54f-f548-4395-83ad-6a7ecc68b03c · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

Visual Lexicon: Rich Image Features in Language Space PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.983804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.983804Z digest=sha256:f3e9d90c010a025fa8fb2498c528c70dbca7e3e0a99ae8732e43695affaa452f

Observation 5ed146b8-028f-4577-a6e2-05d74d7da021 · outbound

This paper cites Deconstructing Denoising Diffusion Models for Self-Supervised Learning.

Visual Lexicon: Rich Image Features in Language Space Deconstructing Denoising Diffusion Models for Self-Supervised Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.993646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.993646Z digest=sha256:0c72ffeb1a397281cb5a5b9d6936b0ce0d05aaf77da258d54e8d13c8f38419a6

Observation 7541b5ae-c73c-49bb-853f-026e0c757138 · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

Visual Lexicon: Rich Image Features in Language Space Reproducible scal- ing laws for contrastive language-image learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.003755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.003755Z digest=sha256:75e040ba97336997ceb2dc971d0903a819d6a5d755687a3bb1eddbe5414b3713

Observation 5f2108ee-e53b-4aaa-9503-3a674420d1c9 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Visual Lexicon: Rich Image Features in Language Space Imagenet: A large-scale hierarchical image database

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.010841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.010841Z digest=sha256:816ba6b901c6f6b1d4ee1116e5bc61fd3bd22596add4b66ff7042bb42a81930a

Observation 03c98d72-164b-46e9-8f91-76a18d6c12ef · outbound

This paper cites Large scale adversarial representation learning.

Visual Lexicon: Rich Image Features in Language Space Large scale adversarial representation learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.017006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.017006Z digest=sha256:ff71b5ab7446e1f5984ea16c03f974ce8ba40cde1f8683e07270912cecf44742

Observation dfbf0931-9762-45ec-a0db-b6e1c13dcffa · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Visual Lexicon: Rich Image Features in Language Space An image is worth 16x16 words: Transformers for image recognition at scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.023580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.023580Z digest=sha256:5dd103d35f3076e9dd78149ffcba4b315c3494a35224ec2be37363a0b90c0f8e

Observation 6ccb5774-6814-46c8-833f-2be2bdda1bf9 · outbound

This paper cites A new algorithm for data compression.

Visual Lexicon: Rich Image Features in Language Space A new algorithm for data compression

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.410098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.029405Z digest=sha256:87aec734019678691683ff70269427c28f90a53e3a68f35ffe4ad24ec8e484f0

Observation af91a1d3-e1e5-4ffb-918d-634b705575b9 · outbound

This paper cites An image is worth one word: Personalizing text-to-image gen- eration using textual inversion.

Visual Lexicon: Rich Image Features in Language Space An image is worth one word: Personalizing text-to-image gen- eration using textual inversion

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.383161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.034815Z digest=sha256:4ef723445e19bf9e4c3b8ed8354556d8a934abac8a95e47cd2b9d4fb2af461d8

Observation 1d414d07-28c6-4f8d-89ae-da66703c8d72 · outbound

This paper cites Instructcv: Instruction- tuned text-to-image diffusion models as vision generalists.

Visual Lexicon: Rich Image Features in Language Space Instructcv: Instruction- tuned text-to-image diffusion models as vision generalists

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.352919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.046351Z digest=sha256:f2bd25e8d5d12e872aec09290c315c3521d0662749d4b431930e340b47afbcb4

Observation 9ebe5898-7155-478e-9f7c-984e83ac06bf · outbound

This paper cites Imagebind: One embedding space to bind them all.

Visual Lexicon: Rich Image Features in Language Space Imagebind: One embedding space to bind them all

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.056655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.056655Z digest=sha256:3b98f26c8ca734c6737ca94e3d9e8346bb67c38768887113c0134bf37d525edd

Observation 8fcdbec4-24b4-4a5d-9cd8-4672cdd60922 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Visual Lexicon: Rich Image Features in Language Space Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.313708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.065525Z digest=sha256:6779f91689d2487a344752314dfd266e4888995ce01b45b22234ae1c4d74e716

Observation 3d9bdb7a-b3b2-4aba-8c93-62428a98577e · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Visual Lexicon: Rich Image Features in Language Space Vizwiz grand challenge: Answering visual questions from blind people

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.072741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.072741Z digest=sha256:ebaa5d3b2cc275b24343ad21f8cb85073753acfe25abdeb02cadc6b0a32dcefd

Observation dfda47b2-9d57-4352-bc29-30acf26db66c · outbound

This paper cites Masked autoencoders are scalable vision learners.

Visual Lexicon: Rich Image Features in Language Space Masked autoencoders are scalable vision learners

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.281433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.078927Z digest=sha256:6ee872ae78a014d8af184b288c5946a3cc6f58080dbdded771eaed80b12271f7

Observation 7bc44f81-38ce-4805-bc07-4364fc897ec9 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Visual Lexicon: Rich Image Features in Language Space Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.089100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.089100Z digest=sha256:149a42c5c4d8c7a39013d0abfc8326be8d61aa12c1b0d1844886b746e2b3c591

Observation ec23e6a0-1c94-4aab-a2a6-dc8159441e07 · outbound

This paper cites Autoencoders, mini- mum description length and helmholtz free energy.Advances in neural information processing systems, 6, 1993.

Visual Lexicon: Rich Image Features in Language Space Autoencoders, mini- mum description length and helmholtz free energy.Advances in neural information processing systems, 6, 1993

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.250336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.098671Z digest=sha256:b5207eb290a3ca837639bc94b51d34e9383babfc5acd009483fecefc698009aa

Observation f8c2c37b-6659-46f5-b568-2aeafbbc2a9a · outbound

This paper cites Classifier-free diffusion guidance.

Visual Lexicon: Rich Image Features in Language Space Classifier-free diffusion guidance

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.108679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.108679Z digest=sha256:bf8802d587d85800ca3fe1250b7ddb343cf070951fd3e893eb6134f74a1ec4af

Observation 38770467-6fb6-47fe-a205-3d69acc679aa · outbound

This paper cites Denoising dif- fusion probabilistic models.

Visual Lexicon: Rich Image Features in Language Space Denoising dif- fusion probabilistic models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.115996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.115996Z digest=sha256:9b17a5d4b377d931208277a6a9587a7b5139a6358666e11ccadbff1c9df8d6dd

Observation bd3f989b-bf23-49db-b91e-65274e5f4643 · outbound

This paper cites SciCap: Generating Captions for Scientific Figures.

Visual Lexicon: Rich Image Features in Language Space SciCap: Generating Captions for Scientific Figures

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.125210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.125210Z digest=sha256:8ad0f4eda1915330f29767116359b902e7a3fc731b277a52e8e34054aa193f37

Observation 77acda40-70fd-45e1-8a62-9ffb034197ec · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Visual Lexicon: Rich Image Features in Language Space LoRA: Low-Rank Adaptation of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.132713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.132713Z digest=sha256:655996d0c6712eb219f0d8e171892448c2456295ffa130237deaf291b6595060

Observation 820a3d38-192c-4903-8b28-2b68487e3ded · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Visual Lexicon: Rich Image Features in Language Space Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.205052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.140234Z digest=sha256:208fec5a6a8127ee54d2d27cde66004cfd035cef83373df1b5b9508061a450fb

Observation 214ad2ed-739b-401b-a78e-4a7e7ceb5abb · outbound

This paper cites Soda: Bottleneck diffusion models for representation learning.

Visual Lexicon: Rich Image Features in Language Space Soda: Bottleneck diffusion models for representation learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.187455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.153036Z digest=sha256:3aa31a27ec29d5cc73c029a2ca175ee5c7b1d7f144482807f8f1029ea489cb0d

Observation d048583f-2d3a-4248-995e-33dbe42e710a · outbound

This paper cites Open- clip, 2021.

Visual Lexicon: Rich Image Features in Language Space Open- clip, 2021

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.168440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.163968Z digest=sha256:0a25481286398d1a4dd7fa097ee069d85878eedb02e1817bb267a35130effa8c

Observation cbb5b517-af0f-4458-aca6-fa74ef836242 · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

Visual Lexicon: Rich Image Features in Language Space Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.145340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.171855Z digest=sha256:7a14a87130aed419a6334692b1cd410ff5e4d128164d835894c387a0d106c7f7

Observation 48d02701-c186-4f6d-a14f-6b27347adc43 · outbound

This paper cites Auto-Encoding Variational Bayes.

Visual Lexicon: Rich Image Features in Language Space Auto-Encoding Variational Bayes

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.180039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.180039Z digest=sha256:1a43ac0dccc437cea5ce78b71f70d09b8e47153cbd3c3b32edca105139af8e5a

Observation 5722e5a2-58b1-469f-8f16-4016ff76d8ee · outbound

This paper cites Your diffusion model is secretly a zero-shot classifier.

Visual Lexicon: Rich Image Features in Language Space Your diffusion model is secretly a zero-shot classifier

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.189781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.189781Z digest=sha256:dc64fdfd5bbbdc9cfe037d812dac36da4af1e1d8651f9cb86efb4777613ab5d6

Observation 60760f2a-1876-4691-81ba-2bfdefc76522 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Visual Lexicon: Rich Image Features in Language Space LLaVA-OneVision: Easy Visual Task Transfer

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.196482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.196482Z digest=sha256:8c648068bd1a65728927338cedb7bc85291495a93aa38a13f3b4fffa2da8c1e1

Observation 0f506868-c561-40dd-ba55-39aa22e4f875 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation.

Visual Lexicon: Rich Image Features in Language Space Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.108917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.203918Z digest=sha256:dda264ffa5c6778abf3ab59e0f9ce50496b1139a85933a7425538bcf70f2a539

Observation 2881563e-06ee-46ce-a4d5-84293a2247eb · outbound

This paper cites ImageFolder: Autoregressive Image Generation with Folded Tokens.

Visual Lexicon: Rich Image Features in Language Space ImageFolder: Autoregressive Image Generation with Folded Tokens

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.217594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.217594Z digest=sha256:d95ee425a66aea6ad4dbd6f8bdd8916f69eb5274729325af6dea253935a10f06

Observation 4234095f-c9a6-412e-920f-8b00f32cbcad · outbound

This paper cites Vila: On pre-training for visual language models.

Visual Lexicon: Rich Image Features in Language Space Vila: On pre-training for visual language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.090493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.226402Z digest=sha256:7e99543996fbebc4457ec91f6b46bd3296cf48e27ad89f7081a3be67c40efc28

Observation 51f159c9-6975-4607-9083-ce26e41d3649 · outbound

This paper cites Microsoft coco: Common objects in context.

Visual Lexicon: Rich Image Features in Language Space Microsoft coco: Common objects in context

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.071696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.231979Z digest=sha256:e75960dc1480d234c7c8769a69c839e7d70acf6ff6f7c03a096afc475eb3103a

Observation 31ef5b21-af4d-4adb-af7c-7cf765f11b39 · outbound

This paper cites Improved baselines with visual instruction tuning.

Visual Lexicon: Rich Image Features in Language Space Improved baselines with visual instruction tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.237076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.237076Z digest=sha256:1fd4092fdb0de181e720c84910a4c4f4ce51d028c122f93ecb153132c64b6ea8

Observation 73a78549-8fda-415f-b5fd-31b0fe850442 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Visual Lexicon: Rich Image Features in Language Space Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.038651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.244513Z digest=sha256:99a63124887534582061e50b18f1d75e984d10c9b5c23815d969eb2c0c590fb7

Observation 16f3383e-3957-4fb6-87bd-f7260ef7f4b9 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Visual Lexicon: Rich Image Features in Language Space Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.251306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.251306Z digest=sha256:268350e284330ebabad6e4e4ab500eb769a8a789fd3699234835889aead30545

Observation e9a4194e-ccde-4b59-b7a6-20b105456862 · outbound

This paper cites Prompting hard or hardly prompting: Prompt inversion for text-to-image diffusion models.

Visual Lexicon: Rich Image Features in Language Space Prompting hard or hardly prompting: Prompt inversion for text-to-image diffusion models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.004154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.256987Z digest=sha256:1b5d80b214384a546830cd5ac88f0ae329d144fdc3b9fd2e4df87d7eeee456c3

Observation 2c298753-c6e2-42fc-aa86-fc2884543f2c · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

Visual Lexicon: Rich Image Features in Language Space Generation and comprehension of unambiguous object descriptions

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.985685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.265281Z digest=sha256:df1d3e9166d45773c191849cc8d137037615ac9dbef473de7247345893efaac9

Observation 11192204-bfe0-41cd-99de-3e2e4013ef88 · outbound

This paper cites Ok-vqa: A visual question answering 10 benchmark requiring external knowledge.

Visual Lexicon: Rich Image Features in Language Space Ok-vqa: A visual question answering 10 benchmark requiring external knowledge

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.966163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.273236Z digest=sha256:6d4637e0aa4f009a1d2a08e4c92a75c9c5449fa86b0e0a30ed5a0b7ba7e22968

Observation 343df608-9cd4-4d54-8b63-dabefd91fb01 · outbound

This paper cites Finite scalar quantization: Vq-vae made simple.

Visual Lexicon: Rich Image Features in Language Space Finite scalar quantization: Vq-vae made simple

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.947678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.282252Z digest=sha256:a081b0541b06eea4b8a4ea63a07296f9d566fc66f94e47d90edc14b21825af74

Observation 7e9a96d0-4db1-4a8f-85d1-d0e9c6846b71 · outbound

This paper cites Null-text inversion for editing real im- ages using guided diffusion models.

Visual Lexicon: Rich Image Features in Language Space Null-text inversion for editing real im- ages using guided diffusion models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.288368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.288368Z digest=sha256:f68a1eda379246f7c41986062aa0a51d453cb03fad5065947b5fc272a70b0e6e

Observation 25210b1e-3cda-4333-a301-8ae368bff820 · outbound

This paper cites Improved denoising diffusion probabilistic models.

Visual Lexicon: Rich Image Features in Language Space Improved denoising diffusion probabilistic models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.917686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.297128Z digest=sha256:ee4312cebe203340f9f5b00328d07dc3d98fd4a3f0701ceb7abb59f65d642ef3

Observation c3683f6d-6730-4d9f-b0d9-bcad9a0e70ad · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Visual Lexicon: Rich Image Features in Language Space DINOv2: Learning Robust Visual Features without Supervision

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.305714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.305714Z digest=sha256:82bcb79b8c14f85439394b846750a7bf734019da6bbd74bbdf9c91f1f25e4575

Observation 780337ba-4573-47fb-973f-c07c08df00a6 · outbound

This paper cites Styleclip: Text-driven manipulation of stylegan imagery.

Visual Lexicon: Rich Image Features in Language Space Styleclip: Text-driven manipulation of stylegan imagery

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.314781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.314781Z digest=sha256:8b50ca522b51e3c504d3b2cd08d0ba2004d362dba557789364dcf80528c722ad

Observation 6972a36d-6ad5-4e07-979e-bd5d9eae5668 · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis.

Visual Lexicon: Rich Image Features in Language Space Sdxl: Improving latent diffusion models for high-resolution image synthesis

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.883818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.329481Z digest=sha256:706596201003817d007d6bba2393ad98b464d01e537e9fd424f5dfb393f9ace7

Observation 3526c95c-494c-45a0-9b7d-392445279843 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Visual Lexicon: Rich Image Features in Language Space Learn- ing transferable visual models from natural language super- vision

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.863188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.346679Z digest=sha256:457c0a243f6899b7b4fb444fdd920f9f758d209f56cd1f28fa7211b8217d68cb

Observation 07ccb4c7-7fe7-4226-9664-470eb4c50e37 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Visual Lexicon: Rich Image Features in Language Space Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.839149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.352693Z digest=sha256:8160b102070a283fd83ae1cb9c2a3efa6930d3106c21de91a07151ac16fe4afe

Observation 97de4469-4aca-463b-ac3c-1f848dbc3502 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Visual Lexicon: Rich Image Features in Language Space Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.360946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.360946Z digest=sha256:9660728a755cecf7d63d92a8830f6c168576de244fe71a4a315d753944cba7a7

Observation 08351e34-873d-4549-b142-3807ea059427 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

Visual Lexicon: Rich Image Features in Language Space High-resolution image syn- thesis with latent diffusion models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.819093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.367294Z digest=sha256:81a7b6f41def42ec6565979d74a4aa3b3e99cc490b7854b90a8bf1de3e656711

Observation 1741204b-4a84-47fc-8dc5-cca29e09855f · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

Visual Lexicon: Rich Image Features in Language Space U-net: Convolutional networks for biomedical image segmentation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.799309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.375945Z digest=sha256:9238b3f832a27c5863e827ccff4c12ec3d5eebeba5099868f5be91db85065893

Observation e9180746-7448-4671-93cd-2e7235e8e822 · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

Visual Lexicon: Rich Image Features in Language Space Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.772700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.383169Z digest=sha256:be438a83f54430e2097d95d51f1ae9d6e8201bcf60ebbfb801c27b2301cdcd38

Observation 726ffdd5-ad93-4285-875c-1b79fdb9ee05 · outbound

This paper cites Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models.

Visual Lexicon: Rich Image Features in Language Space Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.751649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.390493Z digest=sha256:d0fb0f65883797a00e72c0b17ac4b5226c3e7ace40a18e2d60af1fc8b3c09ab8

Observation b21586a3-038e-4654-ad79-7db984e8432f · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Visual Lexicon: Rich Image Features in Language Space Photorealistic text-to-image diffusion models with deep language understanding

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.730199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.398079Z digest=sha256:388333fae8a4d96c1bc3175d7fb20a7aa7c2be081c37dc07bdf352e68a8071f4

Observation 9cb61f0b-e9c7-4677-80ee-4c886789ddef · outbound

This paper cites Progressive distillation for fast sampling of diffusion models.

Visual Lexicon: Rich Image Features in Language Space Progressive distillation for fast sampling of diffusion models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.711979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.405021Z digest=sha256:cc6fdc060036d740bc6025d10f5644cb6514eb44293b2312050e523e04615d8f

Observation e1062836-6fdf-430a-b5fe-b1e2ceafd75c · outbound

This paper cites Neural Machine Translation of Rare Words with Subword Units.

Visual Lexicon: Rich Image Features in Language Space Neural Machine Translation of Rare Words with Subword Units

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.415902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.415902Z digest=sha256:2a549f58a5be0befda533cabe98b4a3fa7f1291088fdba34bee3c4fd72cb27a6

Observation dbb4e182-7d1f-47d6-aba4-56cd33b86dd6 · outbound

This paper cites Adafactor: Adaptive learning rates with sublinear memory cost.

Visual Lexicon: Rich Image Features in Language Space Adafactor: Adaptive learning rates with sublinear memory cost

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.690894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.423417Z digest=sha256:282cc85adc9f079bee5fe130553ffca1c02c84c2a74d5108980ef3df86cbdeda

Observation 98c570c5-8c58-4703-9090-7ceb5fd670a4 · outbound

This paper cites Textcaps: a dataset for image caption- ing with reading comprehension.

Visual Lexicon: Rich Image Features in Language Space Textcaps: a dataset for image caption- ing with reading comprehension

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.432155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.432155Z digest=sha256:32ae2a483bf5cc59f876d7b6083135ca850fe021d71395487505066721a5f36c

Observation 6a0cd468-0675-418e-adb7-e63546c0b297 · outbound

This paper cites Towards vqa models that can read.

Visual Lexicon: Rich Image Features in Language Space Towards vqa models that can read

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.442633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.442633Z digest=sha256:ac04b1a46c5edb2e466ed35297353a703bd3ec2a22b311dd7f1196c25c15ee27

Observation 40168739-51f3-494d-82b2-52916c5d5d37 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Visual Lexicon: Rich Image Features in Language Space Deep unsupervised learning using nonequilibrium thermodynamics

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.638259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.455532Z digest=sha256:9ec52623e5b054252f4ac117eb49d0dbdfde4f037bb315219c95a4b91809d4c0

Observation 9d5d67f3-4ce2-4d40-8d65-5b72cb5ddac6 · outbound

This paper cites Generative modeling by esti- mating gradients of the data distribution.

Visual Lexicon: Rich Image Features in Language Space Generative modeling by esti- mating gradients of the data distribution

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.617318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.465049Z digest=sha256:15bd3477536669bfb3c4f53f2eee12d62917000ed3ce8f1768063e97eb6a6baa

Observation c0291128-0b6e-4d13-b777-f912c63b328c · outbound

This paper cites Rethinking the inception archi- tecture for computer vision.

Visual Lexicon: Rich Image Features in Language Space Rethinking the inception archi- tecture for computer vision

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.471866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.471866Z digest=sha256:5b4c6015e71db42bd7571a928d47ee43a4ed3b4e6f6beea0496b14c3df322d65

Observation 92366efc-2d27-4d38-a7f1-4ad0d2a9d062 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Visual Lexicon: Rich Image Features in Language Space Gemma: Open Models Based on Gemini Research and Technology

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.481135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.481135Z digest=sha256:06559016ac804cfdd75876e9b63fb6cd4fcdc6ef27ee476768a3b7a971ef4d6a

Observation 72cfda0f-2a18-454e-90c0-ee95e2dbf7d5 · outbound

This paper cites Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset.

Visual Lexicon: Rich Image Features in Language Space Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.487818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.487818Z digest=sha256:719fcff246b84784a3aaa335dcc19dd78a456217da8dbf7bfb7bd1a08e0f6e9f

Observation b6ef4604-a898-4f20-9585-20a277c9e7ea · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.

Visual Lexicon: Rich Image Features in Language Space Visual autoregressive modeling: Scalable image generation via next-scale prediction

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.579690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.493448Z digest=sha256:89f95ae90e1b6686abd21e5ea9bb09fb09b24ba33fce4b8a6f8e00d5cf878ffc

Observation 3617e7a4-14ec-4c25-9ab0-df3055614c94 · outbound

This paper cites Con- trastive multiview coding.

Visual Lexicon: Rich Image Features in Language Space Con- trastive multiview coding

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.554902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.498748Z digest=sha256:361a91c4c7e6d223cbe14389bb99a45d96098d4bc3e7c0a4b861313832eb9a82

Observation 787f805e-c67a-4a2f-b217-0d5cea682b99 · outbound

This paper cites Extracting and composing robust features with denoising autoencoders.

Visual Lexicon: Rich Image Features in Language Space Extracting and composing robust features with denoising autoencoders

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.535837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.506084Z digest=sha256:f7f6932742ac887bf2d4929df709ebe480ba8231b6df05a0fe3116fcc53ac053

Observation bed01bb0-dc17-4af1-a093-924f36bfa87c · outbound

This paper cites Diffusion Feedback Helps CLIP See Better.

Visual Lexicon: Rich Image Features in Language Space Diffusion Feedback Helps CLIP See Better

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.515775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.515775Z digest=sha256:8548f05bf0cde3a58a3e761dc9e698782d6f8b82a1b9d4314e0cc934ce9e7f9c

Observation 7ca3aeca-eb7a-4e63-b017-670a2fa7dc70 · outbound

This paper cites Unsupervised feature learning by cross-level instance-group discrimina- tion.

Visual Lexicon: Rich Image Features in Language Space Unsupervised feature learning by cross-level instance-group discrimina- tion

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.508284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.525424Z digest=sha256:e9657dfd7c94f851f5777918a2aa9cacd84e7916c5f84b00d7f9550230623c4b

Observation cefa8917-1ddc-48fc-8ebb-d62875f10f97 · outbound

This paper cites De-diffusion makes text a strong cross- modal interface.

Visual Lexicon: Rich Image Features in Language Space De-diffusion makes text a strong cross- modal interface

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.477684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.535074Z digest=sha256:5eae59a98e599ba30ec68a9916bc28731d29aef943aac71aadf8338115141131

Observation d6b9ab6d-99e3-45f6-abac-85d7c6cd926e · outbound

This paper cites Unsupervised feature learning via non-parametric instance discrimination.

Visual Lexicon: Rich Image Features in Language Space Unsupervised feature learning via non-parametric instance discrimination

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.450250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.544008Z digest=sha256:2d6d66ba616ab3de64c8ed2ae157ea699182c35704216949fcc26342426422e4

Observation ea83b458-1ff4-4fb3-aa0f-d4a9ebc34188 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

Visual Lexicon: Rich Image Features in Language Space Msr-vtt: A large video description dataset for bridging video and language

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.427896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.553657Z digest=sha256:1cee05144a9ff6638c25c24ab3882e5b737752e0d552ac2e781c2e6b831c8c3c

Observation c8af8a4d-ca0b-4740-8d1b-0dc86c4cc6d6 · outbound

This paper cites Open-vocabulary panop- tic segmentation with text-to-image diffusion models.

Visual Lexicon: Rich Image Features in Language Space Open-vocabulary panop- tic segmentation with text-to-image diffusion models

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.401306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.559898Z digest=sha256:f3ef9752f4b684c713c0ebdd7a87e3583e1ee67ee365716c67b1ba1fa8086b25

Observation 13b41e26-cd54-491f-a928-9d7d24d32d7a · outbound

This paper cites Diffusion model as repre- sentation learner.

Visual Lexicon: Rich Image Features in Language Space Diffusion model as repre- sentation learner

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.371408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.569386Z digest=sha256:33d33eac88ea9322b7ccf898c573d4d74881167b5f9df8fc350d8a7caf71ed0c

Observation 8ef7244e-c05c-46dc-8b6c-94b0e5886f8e · outbound

This paper cites Vector-quantized image modeling with improved vqgan.

Visual Lexicon: Rich Image Features in Language Space Vector-quantized image modeling with improved vqgan

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.343354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.579225Z digest=sha256:e3fd92b40e7cec0290152cd5280520b9b00b70b58f97d2f00a926249e15e7149

Observation cfc5324d-733c-4bbe-9044-d46b3037695a · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models.

Visual Lexicon: Rich Image Features in Language Space Coca: Contrastive captioners are image-text foundation models

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.320682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.592159Z digest=sha256:b92d92dd954b9951d37f58ae9e069babc1f20c876ccc116f1fcdd73029ba054b

Observation 17770810-7065-4cf0-9f7f-97014c3b3370 · outbound

This paper cites Modeling context in referring expres- sions.

Visual Lexicon: Rich Image Features in Language Space Modeling context in referring expres- sions

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.298202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.598680Z digest=sha256:4dac56820456035d23267958d1fa49f8be20964a998ecfe51e0f7911250f16c7

Observation cf55616b-e5eb-4651-afb4-d0794ba586d7 · outbound

This paper cites Spae: Semantic pyramid autoencoder for multimodal generation with frozen llms.

Visual Lexicon: Rich Image Features in Language Space Spae: Semantic pyramid autoencoder for multimodal generation with frozen llms

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.273396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.605122Z digest=sha256:6be9352eba785bc6e907432faf6ad11440ba322331b730dccbb4aa71fe09cf72

Observation e45783fa-396d-4806-9a3c-bc9e234330df · outbound

This paper cites Language model beats diffusion–tokenizer is key to visual generation.

Visual Lexicon: Rich Image Features in Language Space Language model beats diffusion–tokenizer is key to visual generation

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.612012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.612012Z digest=sha256:b342fd5a365e2881c00d9ccde59cca2e2afd021498195cebbdf784e06fe107e6

Observation fbda5d99-2584-4952-9df4-2b260fc56b0d · outbound

This paper cites An image is worth 32 tokens for reconstruction and generation.

Visual Lexicon: Rich Image Features in Language Space An image is worth 32 tokens for reconstruction and generation

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.232376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.618321Z digest=sha256:195297bc2724b9284b90980dc91bf645923fdbad5b32d43e93079fbb073e013c

Observation 9cfe70ec-22c6-428e-8c94-9b9fd7270e0a · outbound

This paper cites Soundstream: An end- to-end neural audio codec.

Visual Lexicon: Rich Image Features in Language Space Soundstream: An end- to-end neural audio codec

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.207434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.624640Z digest=sha256:e6c246fad208b3a1dcf312ff57a9b479035a777ef5b8eabb57d4acf173fcab4b

Observation 69467f88-3b85-4b8c-ba8d-eb8f8942aaf0 · outbound

This paper cites Sigmoid loss for language image pre-training.

Visual Lexicon: Rich Image Features in Language Space Sigmoid loss for language image pre-training

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.185620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:22:55.629661Z digest=sha256:58a663a9849c764246405a61a07e149111bca2942faddf8b61c71517e2104c41

Observation 2e2c4c98-8c3b-434a-857d-ddd25aef3edd · outbound

This paper cites Image and Video Tokenization with Binary Spherical Quantization.

Visual Lexicon: Rich Image Features in Language Space Image and Video Tokenization with Binary Spherical Quantization

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.635316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.635316Z digest=sha256:f8895c26fc7d0065f2adb88fad0fb018863fb3f16a70167455a89045efb05791

Pith citing papers

Observation e599697c-7bea-411d-b4f2-60dcfa7a302b · inbound

Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models cites this paper.

Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models Visual Lexicon: Rich Image Features in Language Space

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:54:12.941828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:54:09.796839Z digest=sha256:ea84c22fa83ce6615295121642f4e3a4f65fc076417234f480c805a509186290