Pith. sign in

Paper Citation Record · LEDGER

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation

As of 22 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 1 inbound Pith citation observation for arXiv:2509.01341.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.01341 v1

Coverage vector

measured 93 of 93 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:44:48.021916Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-30T22:45:19.070143Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

93 of 93 outbound references displayed

  • verified exact2
  • verified fuzzy47
  • unresolved42
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5fa19d0f-9669-44ec-a21a-64d7e2a13799 · outbound

This paper cites Revisiting im2gps in the deep learning era,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Revisiting im2gps in the deep learning era,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:40.981376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:40.981376Z digest=sha256:99782a52805f6a2ba136f1120dc01d61545255169c9b67cf982a4cce2a5abd26

Observation b76d9477-03af-47d1-962f-07c90233d94a · outbound

This paper cites Using twitter data to monitor natural disaster social dynamics: A recurrent neural net- work approach with word embeddings and kernel density estimation,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Using twitter data to monitor natural disaster social dynamics: A recurrent neural net- work approach with word embeddings and kernel density estimation,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:41.025665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:41.025665Z digest=sha256:ac76c3e572531bd5a1b5bf32c7e49d7a089a0f55f3991e5131c24f514813f004

Observation 024f83b7-67ec-44e3-a84b-5d45531c9e60 · outbound

This paper cites Suwaileh, T.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Suwaileh, T

Reference 3

Resolution
malformed identifier
no resolver link, observed 2026-08-05T12:44:41.125391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:41.125391Z digest=sha256:ff987e6f8594531afe2af18d29475bd1e8535b98997949a84a6b62994393d0eb

Observation a8c7841b-d39d-45a9-9f6c-f9ae8a5d0677 · outbound

This paper cites True lies in geospatial big data: detecting location spoofing in social media,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation True lies in geospatial big data: detecting location spoofing in social media,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:41.181710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:41.181710Z digest=sha256:a2a988b16e18982b9e60a6c30ca2c3b569e154762c4fc358d33856539bce63e0

Observation 16bea4e7-f42d-407a-97b7-b479f2fc9f79 · outbound

This paper cites Geotagging text content with language models and feature mining,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Geotagging text content with language models and feature mining,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:41.277456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:41.277456Z digest=sha256:b1b6dfb05082acff61f7ce6a0ae0220bfaa33a289f3f52d1c02d5ed8f3146505

Observation 771aeff2-9206-43e8-b5fc-a7447a7b8b91 · outbound

This paper cites Geolocalization and navigation by visible light communication to ad- dress automated logistics control,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Geolocalization and navigation by visible light communication to ad- dress automated logistics control,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:41.345060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:41.345060Z digest=sha256:cd974d619b5a3b78b1dac0df2c6d035dcd4c9907a09b07932eba109b5b57a596

Observation b7229520-d602-4f2f-8664-efb844d2a8ec · outbound

This paper cites Ubiquitous real-time geo- spatial localization,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Ubiquitous real-time geo- spatial localization,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:41.419168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:41.419168Z digest=sha256:b9b85e3e3178b3dc7158c8b1f2dfc86a7c602bd15eab466b68893810c54421e1

Observation 88b0bc38-78ea-455d-9970-9d8483262d5c · outbound

This paper cites Single-image localisation using 3d models: Combining hierarchical edge maps and semantic segmentation for domain adap- tation,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Single-image localisation using 3d models: Combining hierarchical edge maps and semantic segmentation for domain adap- tation,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:41.522988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:41.522988Z digest=sha256:9d49051f78a0e2e106e959fa93a7c472f5d365a86854ab527d233aa1c16ee572

Observation 4bbe7932-c350-4647-829b-9ebf811bfc87 · outbound

This paper cites Analysing gender differences in the perceived safety from street view imagery,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Analysing gender differences in the perceived safety from street view imagery,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:41.614335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:41.614335Z digest=sha256:40c8856206cc8490d508a4baa1ed155e50167fa2b4b14319046f55ab85a2a0f9

Observation 3f74abc1-638c-4319-8af3-9e05efa59f49 · outbound

This paper cites Multi-level urban street representation with street-view imagery and hybrid se- mantic graph,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Multi-level urban street representation with street-view imagery and hybrid se- mantic graph,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:41.726845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:41.726845Z digest=sha256:1830a6c39afb2ad2ebf62e7ef9e3cbe73830934a9b2a167faf317570d0b6e3cb

Observation 3506388f-4059-4465-a4bf-e7a20ac0cac1 · outbound

This paper cites Street view imagery in urban analytics and gis: A review,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Street view imagery in urban analytics and gis: A review,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:41.790355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:41.790355Z digest=sha256:1ebb08e5d51434f813df67b0a7e668b9cd66bc817b49c687e60b122b280a584b

Observation 081a131e-4228-4da7-ba8b-f11797a4bbc9 · outbound

This paper cites Global streetscapes – a comprehensive dataset of 10 million street-level images across 688 cities for urban science and analytics,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Global streetscapes – a comprehensive dataset of 10 million street-level images across 688 cities for urban science and analytics,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:41.888055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:41.888055Z digest=sha256:ae1d76a8eebbdb017ba8a4712dca2c425a03b63e7920478cddacee75cc584033

Observation fce76d2a-e0d6-4691-badb-072bdf50e4ad · outbound

This paper cites OpenStreetView-5M: The Many Roads to Global Visual Geolocation.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation OpenStreetView-5M: The Many Roads to Global Visual Geolocation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:44:48.742246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:41.969293Z digest=sha256:73a2c4e002c6da2746503b50b1e088c74c1df97183ab116756511389919fe384

Observation 4808d3e5-590b-4cf0-acf1-8310d6240795 · outbound

This paper cites Are these from the same place? seeing the unseen in cross-view image geo- localization,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Are these from the same place? seeing the unseen in cross-view image geo- localization,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:56.540922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:42.079359Z digest=sha256:f14b5945802f854da5d86ad70c25298243bec8f2337a54c5ca6d7782e8960a6c

Observation b4f6c5bd-aff8-419d-91e7-70419b6007a0 · outbound

This paper cites Geolocation by light: accuracy and precision affected by environmen- tal factors,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Geolocation by light: accuracy and precision affected by environmen- tal factors,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:56.377224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:42.124308Z digest=sha256:ed54f389695f44375e77e58d5bcac8801e0eff9d7e34633d9fed12af4c91d342

Observation 36f95602-e7d4-4427-af8b-057c74f1f65f · outbound

This paper cites Season-invariant gnss-denied visual localization for uavs,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Season-invariant gnss-denied visual localization for uavs,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:56.217342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:42.214980Z digest=sha256:10dc9def507fffd2aa8e8f6088e0f9cf61745a50c840510c71701bebf1a294c6

Observation 43e8d0de-3932-4722-b3f0-0c164c4545df · outbound

This paper cites Language Models are Few-Shot Learners.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Language Models are Few-Shot Learners

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:42.305139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:42.305139Z digest=sha256:9556f5f4db47577448a156e1783d624d926ec97a616f27d293864db616904d14

Observation e79b1727-a3c1-4020-bac0-4b1d97dc7c64 · outbound

This paper cites A survey on multimodal large language models,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation A survey on multimodal large language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:56.082477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:42.358126Z digest=sha256:cf4b424565d666033fb78b01669c66f7f65796696fe49daefb73005a994c95eb

Observation e51b2922-6f19-4764-98cc-882bd478306d · outbound

This paper cites Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:42.489002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:42.489002Z digest=sha256:131e454c60f83b7d585437b6ddf0d5f140f4047675aedf7dbeb67d3a2d2dcfca

Observation 493c2c24-7bf3-498d-8abd-56504bbe5e54 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:42.563018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:42.563018Z digest=sha256:cd51e4bad7919797f8b36f36fb3a5dc202ae838060e6421d919a1754ba7dc3ec

Observation 9430a234-73f3-4c85-bddf-46fa60b38f47 · outbound

This paper cites A review on large language models: Architectures, applications, taxonomies, open issues and challenges,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation A review on large language models: Architectures, applications, taxonomies, open issues and challenges,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:55.874716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:42.666903Z digest=sha256:3179b4425c9e14f433ec1498e16107594a09f11dee9e83d5cf0ba6f8d988c25e

Observation b73bf6fc-580c-496f-89ab-2870deed2639 · outbound

This paper cites GenAI-powered Multi-Agent Paradigm for Smart Urban Mobility: Opportunities and Challenges for Integrating Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) with Intelligent Transportation Systems.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation GenAI-powered Multi-Agent Paradigm for Smart Urban Mobility: Opportunities and Challenges for Integrating Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) with Intelligent Transportation Systems

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:42.707554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:42.707554Z digest=sha256:7f4be655aaa30a02043f96351fe6b84c7c6fdf29817b59fdc3efde4975366406

Observation 6f273de2-1013-4605-9588-b66a2a80c294 · outbound

This paper cites Im2gps: estimating geographic information from a single image,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Im2gps: estimating geographic information from a single image,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:55.670950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:42.772464Z digest=sha256:f9f8481b137c5c20fd3c22bc7937bb425377ed4d4a0ddba6ed656344db5933b5

Observation 173adaf7-efd6-450b-b988-2678e6a4b668 · outbound

This paper cites Planet-photo ge- olocation with convolutional neural networks,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Planet-photo ge- olocation with convolutional neural networks,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:55.530668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:42.841226Z digest=sha256:42784ea6ee633f73154f2acf53a2142a4da15d1094fb03bb17991366661d4ed3

Observation 8cbb97f6-b013-4b58-a99b-ce140c52c0b6 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Learning transferable visual models from natural language supervision,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:42.928094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:42.928094Z digest=sha256:37fe939c4676c43dd507f330ab47f87d06ead2c131b44acb67a65def9bbdc572

Observation 27b08048-bb3a-4248-977f-078764eb2284 · outbound

This paper cites Sigmoid loss for language image pre-training,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Sigmoid loss for language image pre-training,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:55.322009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:43.019903Z digest=sha256:8460f4264b3fcb9277236c81786c9451dc7b616713a62b98c1971f0cc08dc167

Observation 1f1bb282-a6e3-4cfc-aa3b-c25e37926df8 · outbound

This paper cites Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Blip-2: Boot- strapping language-image pre-training with frozen image encoders and large language models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:55.149007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:43.105115Z digest=sha256:688007e68cbf642e40b166aa8101b52106993c97e1b8031199819e36acdeb4bd

Observation 0c04a5b3-2919-48d1-b83b-cb4c7429751f · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation DINOv2: Learning Robust Visual Features without Supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:43.178571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:43.178571Z digest=sha256:c047e540bc6fd2610716ac102e8bab311b0900f37f26e9d6c69cad2bb759bf97

Observation 6d0e8013-5120-4b86-a4bf-7d3e14a16775 · outbound

This paper cites Attention Is All You Need.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Attention Is All You Need

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:43.231966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:43.231966Z digest=sha256:66ebee602aa4365f7c8ec10621d1f96e492370f355171ae96d002bea7417ebd2

Observation de1812bf-cf53-49e2-9d21-5a86bc984f1a · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:43.306692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:43.306692Z digest=sha256:0f34ab36793b8c863862ca63a024d24f2f36e1becec2288edc1d677d8b73abed

Observation 880fe511-f59c-4986-bace-76c22b63dc09 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:43.365779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:43.365779Z digest=sha256:61ee6f038c6cf34e9edcb1e47b4c943caae52a60ab4e35f207a2d783401e4920

Observation 5952b714-538c-462b-8f3c-8d3eb4fae4ee · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation PaLM: Scaling Language Modeling with Pathways

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:43.397239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:43.397239Z digest=sha256:a9c57ee9a843eda10674612fd6920057ee53d85d1a6e4f0c951b7ece9b2cb4bd

Observation 7dfe23a3-ca38-4cfb-8bec-beecad2de322 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation LLaMA: Open and Efficient Foundation Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:43.458143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:43.458143Z digest=sha256:f4a1cbec13b543172a08a2d17605eb48e3faf1a400f21537de9e9101cf7f516b

Observation 7a24f0f9-8d01-4a98-b6d0-332ecb44359e · outbound

This paper cites GPT-4 Technical Report.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation GPT-4 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:43.561992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:43.561992Z digest=sha256:96c17451be2b5bf11bfb0b3e79dbcb8a712f5680191493b5a93b9ebb9efa9710

Observation 56985dd0-76ec-4d9d-994e-99a6486f3d71 · outbound

This paper cites Visual Instruction Tuning.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Visual Instruction Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:43.640545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:43.640545Z digest=sha256:98f3332fc6bd876593744eaeba074cc56c732786277eeb3179d1be75c4fb936c

Observation de09c3ad-7eda-4d7b-a562-b65590640e74 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:43.720036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:43.720036Z digest=sha256:55c45bcc99a122446d512ce73f3e15b954311f9c67b4e367c3ffd3471a79ba92

Observation 8a048d35-ada0-40b1-afb1-687c0dbf2dd4 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:43.811070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:43.811070Z digest=sha256:2cc4a6de5d04c4914b277afaa7d77207889bd60e2e533128b2d20f440f674dbc

Observation 1156148b-7d61-4250-873f-b0c32602c1a8 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:43.921351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:43.921351Z digest=sha256:e532e55e8c8387cfdcd45e16a486f8767546d2db4c3e10c46e84abb03e90bb06

Observation e9c480e1-b2bd-4b84-a56a-c61e657f4c70 · outbound

This paper cites The Llama 3 Herd of Models.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation The Llama 3 Herd of Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:43.967690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:43.967690Z digest=sha256:4bf1feffb1f2811ff1d546d8f9e492710467a72a6ab16f3ed9d73287bd50272e

Observation 8712c688-fe18-4866-a5a4-634a10446085 · outbound

This paper cites Show 11 and tell: A neural image caption generator,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Show 11 and tell: A neural image caption generator,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:55.022694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:44.022186Z digest=sha256:61f1b06d1a33ba580df589bdb50e82f5f896b83df6406b007da3559a9335e9aa

Observation aecaac5b-7311-4b92-83d1-e3c47b25eb28 · outbound

This paper cites Deep visual-semantic align- ments for generating image descriptions,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Deep visual-semantic align- ments for generating image descriptions,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:54.903233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:44.099409Z digest=sha256:ca4f837e98360551d606a7979852cc7a4785b591506f62461499ad609be5c879

Observation 1919ff89-0298-4741-bb0a-01b5ab71554e · outbound

This paper cites Vqa: Visual question answering,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Vqa: Visual question answering,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:54.764109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:44.178128Z digest=sha256:0a674bf29788397fe08efe61167b575e7867a6c9e25a09103aa725fe42845d7e

Observation 5a9f3fb1-e305-4d88-99e3-e2b7ceca5c3c · outbound

This paper cites VSE++: Improving Visual-Semantic Embeddings with Hard Negatives.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation VSE++: Improving Visual-Semantic Embeddings with Hard Negatives

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:44.266231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:44.266231Z digest=sha256:b232e2964b68d55e45eb1c70c62e8b2f48212e9f9c8ed44a5167f60a801d2a2f

Observation f7b583ba-5e7b-4b2b-a552-8b035e747516 · outbound

This paper cites Vilbert: Pre- training task-agnostic visiolinguistic representations for vision-and-language tasks,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Vilbert: Pre- training task-agnostic visiolinguistic representations for vision-and-language tasks,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:54.632746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:44.317093Z digest=sha256:f711063823d073c1d5ab567ba265e2ae0d78df284b6571472dd1aad0a9ee34fe

Observation f8e989bc-bb5c-409d-9c7e-74c6df9fbc68 · outbound

This paper cites Img2loc: Revisiting image geolocaliza- tion using multi-modality foundation models and image- based retrieval-augmented generation,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Img2loc: Revisiting image geolocaliza- tion using multi-modality foundation models and image- based retrieval-augmented generation,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:54.398137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:44.365095Z digest=sha256:50242091695b041ac53349fd68ceefe08add170f358672dcd0989c2bc0d3ee87

Observation 93247d2b-4068-4797-a65f-f030b1440fa5 · outbound

This paper cites Large-scale image geolocal- ization,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Large-scale image geolocal- ization,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:54.136973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:44.452782Z digest=sha256:a1bfba0c1bfef5a5050b9fb130e55c3cd8dcba81d8370a0367764d0874b47c1e

Observation 8fdf0d41-304f-4887-ae07-8ccb283d5ad7 · outbound

This paper cites What makes paris look like paris?.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation What makes paris look like paris?

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:53.914623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:44.569135Z digest=sha256:ecad71b522bfd542876f6dcd2cb0be8b7c349946dea439513fd4ff0b2d15d05d

Observation ec9d0991-4e47-4dd2-8ef8-afd231c9a84c · outbound

This paper cites Ge- olocation estimation of photos using a hierarchical model and scene classification,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Ge- olocation estimation of photos using a hierarchical model and scene classification,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:53.666246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:44.628214Z digest=sha256:d2f2dbc03049f2e637d509cbf96afb41b3aad02cd7aafd981bf33e3ba1943d91

Observation 2c97a1c7-eadb-4ea7-9a86-3a98694af1f2 · outbound

This paper cites Cplanet: Enhancing image geolocalization by combinatorial par- titioning of maps,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Cplanet: Enhancing image geolocalization by combinatorial par- titioning of maps,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:53.458383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:44.703508Z digest=sha256:9ee0953b5b0b40a3ccf9d68805159e0c12dd91a796f108f21fb8caf50b55b84e

Observation 390b479b-8aca-4e18-8c60-f7c1dfd23d13 · outbound

This paper cites Where in the world is this image? transformer-based geo-localization in the wild,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Where in the world is this image? transformer-based geo-localization in the wild,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:53.316493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:44.778255Z digest=sha256:03df5e7f2f63eca77f7ed590c5dbc2da960753a59720bb6cd9abc1ff5ee1b964

Observation 11958203-564f-4fc1-8d4b-45a08ff3e5b8 · outbound

This paper cites Where we are and what we’re looking at: Query based worldwide image geo-localization using hierarchies and scenes,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Where we are and what we’re looking at: Query based worldwide image geo-localization using hierarchies and scenes,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:53.130919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:44.833945Z digest=sha256:f93526671f8c9344347487dc2d2682eeaca3f259f9530dd94bb588de1dd4aaed

Observation aa5af6f2-8168-400d-9857-daf113528395 · outbound

This paper cites Transgeo: Transformer is all you need for cross-view image geo-localization,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Transgeo: Transformer is all you need for cross-view image geo-localization,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:52.962477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:44.909485Z digest=sha256:1075a80192600a38585578c196e7a3508925150168bd3dea9ada4de2cec8cfe1

Observation d7573abb-40f9-48b6-b838-05d718823ca5 · outbound

This paper cites Joint representation learning and keypoint detection for cross-view geo-localization,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Joint representation learning and keypoint detection for cross-view geo-localization,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:52.830862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:44.934565Z digest=sha256:2fc2e06dd242345889ecd3ed9dd3f4315fbef44b2380441bf03a10f897771c87

Observation bdeecf7a-b745-4f32-b58b-6bac9ee8931b · outbound

This paper cites Cross-view geo-localization via learning disentangled geometric layout correspondence,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Cross-view geo-localization via learning disentangled geometric layout correspondence,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:52.674128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:44.984337Z digest=sha256:cefe564d0612ccfc15dd719cbceaef61b8f9827119d3cc3fb6f426b8512cbd63

Observation bf70c786-b281-4264-959d-8b9a991d6abd · outbound

This paper cites The benchmarking initiative for multimedia evaluation: Mediaeval 2016,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation The benchmarking initiative for multimedia evaluation: Mediaeval 2016,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:52.524673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:45.045779Z digest=sha256:65f42b9c728b3495e272d4ded7dff80aefdd040f6d7f8bc6dc0f4a6606cc5a9c

Observation f1089739-13c9-429e-b772-dbdf4d26438f · outbound

This paper cites Inter- pretable semantic photo geolocation,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Inter- pretable semantic photo geolocation,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:52.385187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:45.115249Z digest=sha256:a0c44a1db958d931d531938b95a437fc3232e280fbe1e594a2edc60fc827bd53

Observation 9c300b8b-4080-48a5-a405-fab3e6ac5a58 · outbound

This paper cites The faiss library,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation The faiss library,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:45.180435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:45.180435Z digest=sha256:a1390151a8a9114c7fa5fed408396e3c497fcde879bb8ca9f2c39034cd13efae

Observation 6b59d702-478d-4748-9765-736baf23ce6f · outbound

This paper cites Billion-scale similarity search with gpus,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Billion-scale similarity search with gpus,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:52.215129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:45.306594Z digest=sha256:9ea9a38276dbdd2e7f7ebe4c5f2178ec99a4bad73eb6cb231ea54f02597da9ff

Observation d07ba2f8-7eb1-4b40-bef2-faa9cf972975 · outbound

This paper cites Yfcc100m: The new data in multimedia research,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Yfcc100m: The new data in multimedia research,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:52.074636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:45.374822Z digest=sha256:92a98bcc6e39d1aadc9d4766d72750990ecaafb99122ba1b6d554207faa9974a

Observation 959b2052-497d-4fcd-bcb4-4a0ffbe65d30 · outbound

This paper cites Geopositioning accuracy assess- ment of geoeye-1 panchromatic and multispectral im- agery,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Geopositioning accuracy assess- ment of geoeye-1 panchromatic and multispectral im- agery,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:51.957721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:45.489852Z digest=sha256:4f48d6c99eba38e8631fb7ea5e7a9cbb23261bd0d5d2c2ad81791f333f485725

Observation e993cba4-d829-4ba5-a004-f6081c766d98 · outbound

This paper cites Autonomous smartphone-based wifi positioning system by using access points localization and crowdsourcing,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Autonomous smartphone-based wifi positioning system by using access points localization and crowdsourcing,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:51.803318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:45.575701Z digest=sha256:419731a1c27a403f17f9bda80ba2a60530848edc13b9eff22f8ca1862b06315e

Observation 61f78d01-b827-4d49-9096-c970a99a6bb6 · outbound

This paper cites A faster and more effective cross-view matching method of uav and satellite images for uav geolocalization,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation A faster and more effective cross-view matching method of uav and satellite images for uav geolocalization,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:51.613103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:45.649573Z digest=sha256:9b992d1193c9763789a60f3ee6fccd9178aa1388d94e405e31e963c4820aee11

Observation 8aeead03-afc3-4fad-baf3-d540f0bf5a74 · outbound

This paper cites Efficient localisation using images and open- streetmaps,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Efficient localisation using images and open- streetmaps,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:51.440018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:45.764378Z digest=sha256:df6cd3ae1c0ac53adc0a710a5fb4187e2bb72b2b9d4f60f7c4bc4795da145a2e

Observation 4f5cec11-b144-453e-b553-abd7e0dd097c · outbound

This paper cites Automatic discovery and geotagging of objects from street view imagery,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Automatic discovery and geotagging of objects from street view imagery,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:51.294148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:45.894119Z digest=sha256:5674c4a7ca0a673b141bd60fe90ef4d02f4e19d13d536ee412fd79b18a90e4e6

Observation c3a3bb35-a88f-4089-a5e0-90f4563999ed · outbound

This paper cites Swin transformer: Hierarchical vi- sion transformer using shifted windows,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Swin transformer: Hierarchical vi- sion transformer using shifted windows,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:51.109136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:46.013528Z digest=sha256:2908c7f67baa58290c2b9fcff178e18945ea9205936b229561a5529b711bc1ad

Observation 6263de81-4f9a-4ef4-ba6f-e56112e6c771 · outbound

This paper cites Pigeon: Predicting image geolocations,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Pigeon: Predicting image geolocations,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:50.940559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:46.092102Z digest=sha256:f030e21e49ad99533efb2bdf5ee3b6a4b24031aed496e77cbdc12c4fdc9e8d8a

Observation d03ea6a7-3a1a-4523-849d-a31a1b2b85e6 · outbound

This paper cites Geoclip: Clip-inspired alignment between locations and images for effective worldwide geo-localization,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Geoclip: Clip-inspired alignment between locations and images for effective worldwide geo-localization,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:50.819586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:46.211759Z digest=sha256:551ed3d18e9c3d3c86d952ff46d1e33e4d70888e98db936e934b5e4ce3dcabc2

Observation f7239308-2b12-4de0-a9a4-ed0315e71564 · outbound

This paper cites Openclip,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Openclip,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:46.309752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:46.309752Z digest=sha256:d61b245b1be403f896d387b99af25b331c1f7cf4c0edecc5d8dfa553daf66aa6

Observation dc2a882b-f5af-4897-8be7-d3fe031030d6 · outbound

This paper cites Learning generalized zero-shot learners for open-domain image geolocaliza- tion,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Learning generalized zero-shot learners for open-domain image geolocaliza- tion,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:50.702106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:46.413062Z digest=sha256:3adb1e7633326c2f0a9b99998075c35d716f2ec3900d74ea2f665dc308e38ebc

Observation dcbbb522-3ff9-461f-ba18-988243d79354 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:46.492478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:46.492478Z digest=sha256:8f8e10c2bac805d1658504d11bddd24645c36b413edc0b151bc95d9abcb22b92

Observation 1da6a83b-226e-4666-88dc-2d8016bebf4a · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:46.547977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:46.547977Z digest=sha256:d69486937ee10ae35a1e420e4a3ebfda1b3560d7da298362ec12d6d345910185

Observation f76c6ad0-07e8-41bb-8f08-e68aeee0b328 · outbound

This paper cites Pixtral 12B.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Pixtral 12B

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:46.631322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:46.631322Z digest=sha256:6d4806cac4c890093f1fa15c26ddfe51edacccf4fa6b97bf27ddade1e2f5b9bc

Observation 748abe4c-ecc3-4ce2-9b2c-bce63cc6da9b · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:46.702453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:46.702453Z digest=sha256:377e7d1a5e232257e2abb5cfee2623b88b33afcd2ddf1f838daed048ca484740

Observation d7bc1fa8-78b5-4306-97a3-7c91d1fc4885 · outbound

This paper cites Hugging Face,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Hugging Face,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:50.551802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:46.773560Z digest=sha256:feeffe6ee9ab0ad9a887c4857bf8c07b13615a05a368444bb84a0e965d8e97f1

Observation 6d1b235c-0761-47ee-8af3-9719f156e79b · outbound

This paper cites A Survey on Evaluation of Multimodal Large Language Models.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation A Survey on Evaluation of Multimodal Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:46.842259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:46.842259Z digest=sha256:373523314a51db38949dfbd7c9c2dca71065f02ec5d7ac07b4832c7082db60d2

Observation 9f9b6e5e-2f7f-489f-be1b-af083429b68e · outbound

This paper cites M3exam: A multilingual, multimodal, multilevel bench- mark for examining large language models,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation M3exam: A multilingual, multimodal, multilevel bench- mark for examining large language models,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:50.430619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:46.926992Z digest=sha256:884278565338a716ecd48a59787368b5889ebe492d2c32c4dad52779668ac49e

Observation a9fe15fb-3c3f-4370-9789-cd2c7f631c8e · outbound

This paper cites Lvlm-ehub: A comprehensive evaluation benchmark for large vision- language models,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Lvlm-ehub: A comprehensive evaluation benchmark for large vision- language models,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:50.276940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:47.011481Z digest=sha256:0ea999d6acdc0450316665a23bb6642e5329d5bc80cb157970128199e72a5e8a

Observation d7caa79b-4a75-4fcf-83cd-f4e60c6b36c5 · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Seed-bench: Benchmarking multimodal large language models,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:50.105747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:47.090861Z digest=sha256:01ab129f514c36ca01cab19140ddd9cb6be14c2ce4fd4e5162fd136ba32fa3b7

Observation 67d6ef8a-69c1-45b1-bcef-c86b5667d0fa · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:47.141517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:47.141517Z digest=sha256:ea7c1ff262cee4ff7b1afa43305579ad9755a04cfe06edd35ad874c417a0240e

Observation c42a2240-f37c-4717-9db7-cab6652f3bba · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:47.201600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:47.201600Z digest=sha256:6201c7730bcbe33d63166559db46c77014d2e85070b7bbbf7dc3c723bfdf189c

Observation 8e712d0c-016b-4fd6-97d5-71b4cbd335f7 · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:47.296569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:47.296569Z digest=sha256:ebb33a9c0544a406fd4f6fbd83ac321db3bfdd6d7f200465f736b8d783aa4dbc

Observation 1098a94b-6014-4063-af94-7e7fcfeb2975 · outbound

This paper cites python-pillow/pillow: 11.0.0,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation python-pillow/pillow: 11.0.0,

Reference 82

Resolution
verified exact
doi, observed 2026-08-05T12:44:48.175217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:47.383043Z digest=sha256:975539b42199c22360116e83a5b0bc5e93f837033506007cfe8d655d87651098

Observation 751105cc-5dd3-4d25-bef7-a8c1233ea7bd · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:47.435962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:47.435962Z digest=sha256:4ba38aac7c3993f11ec3917da8e9dbabe86ebe384d158e00476166526cdd1a66

Observation f3dcd693-9f4b-47d5-b47b-663877724f4a · outbound

This paper cites pandas-dev/pandas: Pandas,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation pandas-dev/pandas: Pandas,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:47.493064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:47.493064Z digest=sha256:671238183b72614c0a980135bb49dda89052011606de1e8b0d0c91c5429b7ca5

Observation 85369fd2-16fa-4b1e-92bc-c239d42da949 · outbound

This paper cites Efficient memory management for large language model serv- ing with pagedattention,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Efficient memory management for large language model serv- ing with pagedattention,

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:49.914094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:47.569033Z digest=sha256:15e0632d8147cd280a265a0f8bcca98069c7a54be064fa4522b9353caea3cfbc

Observation 0eedb2b4-04ce-4b73-b53b-a79dbc5d5de8 · outbound

This paper cites Lmdeploy: A toolkit for compress- ing, deploying, and serving llm,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Lmdeploy: A toolkit for compress- ing, deploying, and serving llm,

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:49.714858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:47.669560Z digest=sha256:e306766f4cc0a54c88dd99f778996bfb620040944a0b20c821d555764a914931

Observation f9858eb9-10ad-4877-8281-99f34e16d669 · outbound

This paper cites Openai api reference - create chat com- pletion,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Openai api reference - create chat com- pletion,

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:49.504315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:47.765929Z digest=sha256:cfed9cbc4ec97c1ea6683745a53859e3c058caf65a187567d21030a9cafff47e

Observation 311f3814-ee22-4dbb-88d1-794690adb77f · outbound

This paper cites Api documentation - sampling parameters,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Api documentation - sampling parameters,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:49.383932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:47.844019Z digest=sha256:0f6f54f510753c0965a46d256b92ff0bd6b28cc28bdeed2b08601bb83aa0a61b

Observation 5038e851-277a-448e-9b02-124b67325bd8 · outbound

This paper cites Geopy - geocoding library for python,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Geopy - geocoding library for python,

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:49.263319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:47.908803Z digest=sha256:d7724d4ec775e8735ff5463f2a61ffa1b04bfd843e9fdac60ae3114137bc56a2

Observation 8b4b807c-fa9b-4e93-acda-433c080f0f2a · outbound

This paper cites Georeasoner: Geo-localization with reasoning in street views using a large vision-language model,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Georeasoner: Geo-localization with reasoning in street views using a large vision-language model,

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:49.086279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:47.968111Z digest=sha256:ad660682ef384fe6e2fa04d5b34631eb31d3794960982669598788790a8eec9a

Observation caab90b3-e5d4-4210-b7f7-ae246b814f48 · outbound

This paper cites Mixed land use measurement and mapping with street view images and spatial context-aware prompts via zero-shot multi- modal learning,.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Mixed land use measurement and mapping with street view images and spatial context-aware prompts via zero-shot multi- modal learning,

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:44:48.834967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:44:48.021916Z digest=sha256:20a32a9d03c3b323659908916a98085b985468c77c08cdf1ba4ebbfb78db1278

Observation d71bd453-1c6a-473e-a8bb-beba2eeef0d7 · outbound

This paper cites Available: https://www.sciencedirect.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Available: https://www.sciencedirect

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T12:44:41.556636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:41.556636Z digest=sha256:8f2024f616177242b225095d2e8207237ed2d0900c1050512ea1c34ff53102b8

Observation b35bafec-c252-40b3-80dc-c3cb4e79a603 · outbound

This paper cites Available: http://dx.doi.org/10.1093/nsr/ nwae403.

Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation Available: http://dx.doi.org/10.1093/nsr/ nwae403

Reference 2024

Resolution
malformed identifier
no resolver link, observed 2026-08-05T12:44:42.417808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:44:42.417808Z digest=sha256:b74b89afdd87cb22b2f6c4f98a5912a843596fb853738c6848ed746fbe4ad1ea

Pith citing papers

Observation a3136a66-8a2a-4914-8b06-1b4b3ce1a573 · inbound

DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization cites this paper.

DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-30T22:45:19.070143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T22:45:19.070143Z digest=sha256:a245a9efed10ed0900dfc437e9f8e6bdcc2aab0ec9402d26014075b06dcb5318