Pith. sign in

Paper Citation Record · LEDGER

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders

As of 20 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2607.06445.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06445 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-08T05:49:25.572256Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact13
  • verified fuzzy4
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch12

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aad35dad-f979-4e0c-89a9-72819d3edc08 · outbound

This paper cites Qwen2.5-VL Technical Report.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Qwen2.5-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:54:33.573554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:b2df87b27c0b34877e4db6a60a76ccbbdbdcf7b80d754ad17fdd3e1b9a095228

Observation 0e49c2a6-39cb-408a-ab15-2dfd325304d4 · outbound

This paper cites Dar, G., Geva, M., Gupta, A., and Berant, J.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Dar, G., Geva, M., Gupta, A., and Berant, J

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-08T05:54:33.547907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:ea2b48377446927b3094e4d1a252a991de78624fefa498dd23e89c1934e7b886

Observation 51cb6c9f-4942-40ed-a626-e84896284fc5 · outbound

This paper cites Analyzing Transformers in Embedding Space.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Analyzing Transformers in Embedding Space

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:54:33.545444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:2acc59443fca792a91425255a807937f4540503d57f07f6bf1a46a363ac66998

Observation 2d884595-620a-4dae-a9f4-eab353ce17d5 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:54:33.551545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:0a10bbbdebe68b52e470e8b609e0f09646b006593dadbc522f83341d550c765e

Observation 6a6d1ddc-11f7-4971-b729-9a09a6eb69e9 · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Transformer Feed-Forward Layers Are Key-Value Memories

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.542671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:61b67597ceaebe446bd880526032f727cace62f31fd1ad66664a0418a5d45803

Observation c1d40b2a-e127-4743-b13c-8b960016e953 · outbound

This paper cites Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.557100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:c0cade8583c453e2f4b42cfe57458853c2aa176dae3f498274fbeff3626c23f9

Observation 153eacbd-34eb-4737-9076-67a803745d13 · outbound

This paper cites Generating an image from 1,000 words: Enhancing text-to-image with structured captions.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Generating an image from 1,000 words: Enhancing text-to-image with structured captions

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-08T05:54:33.569864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:cc13b46f8e5abd4e137b8db59b56a1f7054ef41a45740c89b5dae04ecbc9b428

Observation 57ddb86a-554b-47ca-8298-635aba06c398 · outbound

This paper cites LTX-2: Efficient Joint Audio-Visual Foundation Model.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders LTX-2: Efficient Joint Audio-Visual Foundation Model

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.539837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:83b5fed8279d5a65beab97d6d53128bcacb49e63f4e70c997c357f37b8543237

Observation 86cf63ad-536e-4fb5-86c9-285f1322b384 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:54:33.568137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:4ba39552adb63de433f6e723099592d6e7e5c5b426081fd0fa10f27b9742c9b7

Observation a7f0eec0-b394-48cd-8606-a5ede115aa8a · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders LoRA: Low-Rank Adaptation of Large Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:54:33.566698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:6c3806f6e043717773e19e6d22fb121b5b642c27f642aea9de75a9e2a2f2ef03

Observation 67d74a6e-4e64-4164-a191-9bcd82c72c8b · outbound

This paper cites Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.562879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:60b9eed8c46ab80155acdce4191480a744c88ae72f34d77a758067c63b1c1a3f

Observation f099b476-2c82-4664-bae0-8de9a7610a42 · outbound

This paper cites What's in the Image? A Deep-Dive into the Vision of Vision Language Models.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders What's in the Image? A Deep-Dive into the Vision of Vision Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:54:33.577235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:44beb1b7729fcbf553285a68bfcdb041e2b9cd18b95cc1ca623c85ef931da612

Observation 1a2f12fb-9cd9-43d0-af2f-1dc30bcc99ab · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.547841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:20d6342dd54e62ec7257f04e8ccca53088a0f8734c775d7f0fe756c4560a2d40

Observation dd27a126-1e67-4c4f-acd4-0443c975e361 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.523101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:3e00b543945009c2a969b56b0cc5fab15e09bc584d6bae0aa3e414b332c9dc03

Observation 5a879df2-b80d-49ff-827a-d00aea8bb235 · outbound

This paper cites A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.572461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:69ac525d1d93cb5cea1c4579794ecd097f91989bf6e9db13eb8ac0fc0698cd8d

Observation 36b15485-8046-4abd-9af9-931f173d5ae2 · outbound

This paper cites Evaluating Text-to-Visual Generation with Image-to-Text Generation.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.536126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:f49bc789cc81ffa0c92714ccbe14b62e41e8e090609c08cd15295fc337c5315f

Observation e9ae30cd-31e1-49b4-b55f-ea999ddbf0a9 · outbound

This paper cites Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.575294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:1e1c677f9e8968b3a0679dae152e691ceeba666831b0498d2bc294509d123d6b

Observation f58b044a-45d3-4f44-b54d-0e9b11fbc372 · outbound

This paper cites Seeing but not believing: Probing the disconnect between visual attention and answer correctness in vlms.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Seeing but not believing: Probing the disconnect between visual attention and answer correctness in vlms

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T05:54:33.526094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:5f1a2803631d9e70ed1286b7cdad75d1d6e7fb0baf743f034e1edbbe22e20619

Observation 5125282e-3e2d-40f2-9cec-e3e95d556f27 · outbound

This paper cites Linearly Mapping from Image to Text Space.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Linearly Mapping from Image to Text Space

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.565640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:81ca09bdfaa8bc79c23013ba2947d42de24a55602f740c41f75e482bc2d34fdb

Observation 1e7fcd8f-d830-4ffa-ad08-ca62f0b97579 · outbound

This paper cites Towards Interpreting Visual Information Processing in Vision-Language Models.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:54:33.570640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:f1472351e9446ce9474bb4a1d1b75e325702e163b9dd3988130694c6a20accf5

Observation 481f6402-548c-417a-b779-a5103a20432d · outbound

This paper cites Same task, different circuits: Disentangling modality-specific mechanisms in vlms.arXiv preprint arXiv:2506.09047, 2025a.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Same task, different circuits: Disentangling modality-specific mechanisms in vlms.arXiv preprint arXiv:2506.09047, 2025a

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-08T05:54:33.554573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:383d9e1ab61bfc94d6e3c3380dce79f60aacca70fa64a599f7e24ed94b300475

Observation 80aa4da0-cffb-44f4-9d95-5807e94bcce7 · outbound

This paper cites Rezatofighi, H., Tsoi, N., Gwak, J., Sadeghian, A., Reid, I., and Savarese, S.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Rezatofighi, H., Tsoi, N., Gwak, J., Sadeghian, A., Reid, I., and Savarese, S

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:54:33.883501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:4ae249efdd5a6094a2cd81dafc9335357a99248c71220b342361c440f9517f82

Observation b8c777e6-c106-4ede-88e0-ee5a5a87e18e · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders High-Resolution Image Synthesis with Latent Diffusion Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:54:33.553301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:3f420b7e418ffe36ebc8149d816725bcf7d7af40703c80881e3cc75ad171e02a

Observation 00ecc256-a3cd-4a70-97ea-4ffd0be5259c · outbound

This paper cites Multimodal Few-Shot Learning with Frozen Language Models.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Multimodal Few-Shot Learning with Frozen Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:54:33.580326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:2e3986749fcb80d9bf3a616bb99dac17c5d9f25781a24a37c8b01098892cf098

Observation 7a15bd3c-2bc1-4997-b72a-a2b2311f9306 · outbound

This paper cites OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.521874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:f9b1a93f24ecb8cf74c880d293648d46147cd98a0f2b0298ce6a6648c1cdcdd9

Observation 8311148d-6b60-4cda-b5cc-6498e1eaee1d · outbound

This paper cites The Unreasonable Effectiveness of Deep Features as a Perceptual Metric.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders The Unreasonable Effectiveness of Deep Features as a Perceptual Metric

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:54:33.519294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:53882ddde220ab14226a8390cde1e4a96dc31074294a489888e71e9e37ab0cd1

Observation 9ffb9559-89cb-44ed-96c6-c0135f99bc60 · outbound

This paper cites We choose FIBO due to its strong adherence to spatial layouts.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders We choose FIBO due to its strong adherence to spatial layouts

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:54:33.887144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:fc0bba8399595aae62ecdb1a82bdd4a153e793d7072f264c706ce49112fa1869

Observation 98cc5f86-8a70-4af4-b166-dd5f26f380bf · outbound

This paper cites To stabilize the initial training phase, we apply a warmup period of 100 steps during which the model is optimized using only the L1 loss.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders To stabilize the initial training phase, we apply a warmup period of 100 steps during which the model is optimized using only the L1 loss

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:54:33.885324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:c6d1a51aabec4dd87e4cc2bab0e4182dc4549894475826ca5bc80dc7d1b5e58e

Observation 896f0a33-f442-4b4d-a9f4-1d83dfc38318 · outbound

This paper cites Recolour the bottom right clownfish to be black and white.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Recolour the bottom right clownfish to be black and white

Reference 29

Resolution
malformed identifier
arxiv_id, observed 2026-07-08T05:54:33.563947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:038f06753a8fe63073824627aa5e611b5b4bd5ba56dd2d01a855bd8063fd4f76

Observation e62b7b48-034c-4947-bb26-0b5a997dc9ac · outbound

This paper cites All automated Vision Question Answering evaluations and prompt generations utilizing the Gemini 2.5 Pro API (Gemini Team, Google,.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders All automated Vision Question Answering evaluations and prompt generations utilizing the Gemini 2.5 Pro API (Gemini Team, Google,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:54:33.881647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:795f3d269d90fcdb53f7d9fc7ace80a1a6d8817e011cd0ef8524f886a56b1b30

Pith citing papers

No inbound Pith citation observations are available.