Pith. sign in

Paper Citation Record · LEDGER

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models

As of 7 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2607.16305.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.16305 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T06:33:14.960504Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 784c5693-f5f3-466a-be2e-2504a5d758a8 · outbound

This paper cites STEM: Scaling Transformers with Embedding Modules.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models STEM: Scaling Transformers with Embedding Modules

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.733672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.733672Z digest=sha256:d62fda466debbb0e035e24180a29d71e66edf4bfab83c1ce4595b988c8b8b158

Observation 8c04fa4d-8e33-418b-b576-0591770ed8f3 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:13.908702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:13.908702Z digest=sha256:0b1476a1817d73c5a5ab97020f133e6a281dc2118b5c5ac3f7dd6b2c30ce67ec

Observation d30f5a49-2cbe-4992-96d7-716a8ed7b9b0 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.240999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.240999Z digest=sha256:b64a60f83b217a450ec96cab89d71d48bb4a29a63d62ef5fcf5420e13a2c81a8

Observation 051a2367-bb40-4d66-9543-81b4c9972659 · outbound

This paper cites Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.292946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.292946Z digest=sha256:1f1b0d5ad49caff936cb97fe4315b22ace35d8d905b03fc9c53e2159c1e5db17

Observation 57187ae8-b034-49f7-b94c-ce96b817a597 · outbound

This paper cites Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.485161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.485161Z digest=sha256:34c398a12ffa2f5a7914089f74ce776f4b44a25e7eb81bb7a6fc5a184675987d

Observation 6ff32895-0b63-44fc-a9a8-248689c4b32f · outbound

This paper cites an unresolved cited work.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.625608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.625608Z digest=sha256:a8006e9c651b54e0867700c3d1087f3dae6199c4472f0689502eaa4f60ace630

Observation 54b5273f-5f12-4f84-afcf-746401790819 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.677496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.677496Z digest=sha256:2ddfa0f4d04448631696b4d69d47341c1071f87f7f5c2c2941abce7dcaf07f75

Observation 4598b522-f40d-4cee-a40e-a82abf03f44a · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.864080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.864080Z digest=sha256:93f714f1f257e59f88cc4909d1c2480639ed35d06cc303d3ced659b9dd1cd789

Observation 9c39c865-9dfc-45e6-b333-9540fc28ef60 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models CogVLM: Visual Expert for Pretrained Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.945938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.945938Z digest=sha256:aba05b8207e9567cfa04f94aae5c9e18b64f17d96f30a74bb24c6176c65c37ac

Observation 90a3810e-e305-42f1-ad6f-b6eea24e9090 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.950666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.950666Z digest=sha256:285d76dcfe7e31578a2b16059a3a96fc0fbf1c52cbc4618b91a647a575d27bdf

Observation 5a5680dd-462b-49e5-a284-ee7adb09ce1f · outbound

This paper cites Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.955758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.955758Z digest=sha256:0eca15b98ab795edd99a6a2698052bc9139b3170a7588c5df6e2710c2f5f9cce

Observation e8af6555-174f-49e8-9fb3-4932eeb61e6d · outbound

This paper cites HiMix: Reducing Computational Complexity in Large Vision-Language Models.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models HiMix: Reducing Computational Complexity in Large Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.960504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.960504Z digest=sha256:48d20cf027c2484d5ee9a3b4ed62fbc282890be6acf21f7d337f5b845529843f

Observation 42d5e4bf-7495-4e17-83bf-3cfd89cc6bfc · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.797043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.797043Z digest=sha256:c511bed8647763be8c56e7d601eac9e54c0b7db7fc1cf5ff9625cfb4d766ea3c

Observation 3caceb63-a5fb-44fe-8d15-e6ac4b03e2c4 · outbound

This paper cites Scaling Laws for Neural Language Models.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Scaling Laws for Neural Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.541340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.541340Z digest=sha256:6aa058ffa607b2cbc48221a7c08abbab7871a065b3f2ef79cc1f37cabaf5d034

Observation 968911c8-70be-4eee-8090-d09f98d0d120 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.400943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.400943Z digest=sha256:2d8865e2360911759521b1c22daf40349da61f01315f49926a491487c9e78dc1

Observation 4ec2d70e-32a6-46ad-8d09-a3446753db47 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:13.537859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:13.537859Z digest=sha256:caed8ced6a413881d516ddd7a07746b658220b89f85481c9ea6b1bdf03bafa8e

Observation f39be760-bfdb-4680-ac30-fe4801ecc271 · outbound

This paper cites DeepSeek-V3 Technical Report.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models DeepSeek-V3 Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.128871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.128871Z digest=sha256:4f31700930157afcc48c80ef4cea4aba036d8af2dd12868d1c0372a86eec70cc

Observation 6f5fc2fc-0d84-4509-993b-c3462782ca18 · outbound

This paper cites Qwen2.5-VL Technical Report.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Qwen2.5-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:13.628554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:13.628554Z digest=sha256:9859b16a1229f226502ab5310938b293bee01abe2976a544d98f57cc27f4458b

Observation 33dd7702-4e23-4626-8262-b2c0afafed97 · outbound

This paper cites Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:13.744408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:13.744408Z digest=sha256:0af106e0d9bb3db511c8d7292cfc8ef0c860d442db247924c25b541ba4114d76

Pith citing papers

No inbound Pith citation observations are available.