Pith. sign in

Paper Citation Record · LEDGER

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization

As of 5 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2604.10721.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.10721 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:39:17.229872Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact11
  • verified fuzzy49
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a9bdc644-6d89-4ddc-b444-4432c8f91b75 · outbound

This paper cites Phi-4 Technical Report.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Phi-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:06:03.038475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:492479d0f1fe8bb7684dd5bd6e9639cedede07eb2f1805798b951d8c7392cd05

Observation 58c4810e-ec5d-4931-bbd7-41a215e98f90 · outbound

This paper cites GPT-4 Technical Report.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization GPT-4 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:06:03.164288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:6417946d6473de15416605c7a3d9fec8efddc661865bab366e6c627c3a18209f

Observation 0926454b-3ddc-4ab2-8954-5f1cd00f3a4a · outbound

This paper cites Vision-and-language navigation: In- terpreting visually-grounded navigation instructions in real environments.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Vision-and-language navigation: In- terpreting visually-grounded navigation instructions in real environments

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.882877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:96c31b2c432b054749fdd5421b0f5ce9cb2c262e7357d66176d612c180a58364

Observation 8bf7281f-3a4d-44d2-8a6f-a811aba68607 · outbound

This paper cites Cross-view meets diffusion: Aerial im- age synthesis with geometry and text guidance.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Cross-view meets diffusion: Aerial im- age synthesis with geometry and text guidance

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.876789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:070a8b463ff7dd6289eadc8bd27c8c7a11b50ed100bb443b66da1e0cffc12b37

Observation 1d95bd3d-6d53-462d-ac25-43bdb82aacc6 · outbound

This paper cites LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:03.148741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:5fcf881ebd0ec26ec211ea9d0e24a7e69c73ad0e6a728325b5ac8c1b0fc45f2d

Observation db9cee2a-cf70-432a-8c1d-986b9cbc12b4 · outbound

This paper cites Ground-to-aerial image geo-localization with a hard exemplar reweighting triplet loss.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Ground-to-aerial image geo-localization with a hard exemplar reweighting triplet loss

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.900241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:2e390be3bc0f1b598ac9e1ffb8329dda0a7829037195f1c2f7ececad7a111202

Observation 2e6d8a88-e612-4a8c-b338-11d59d9bf0c4 · outbound

This paper cites End-to- end object detection with transformers.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization End-to- end object detection with transformers

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.871007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:66c44f7f27dc17bc2de570fee4e62855f5f0927303885b29f41f60a34075bc4c

Observation 96e4ba77-2fb2-42e7-a42d-b274cd70c517 · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization A simple framework for contrastive learning of visual representations

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.873906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:2793919649b143ed595e9ce2aa3f9aa7d6af1439a2f802182dd4f6351ab3111c

Observation 79fffef6-e69f-4b3b-bcff-d9d8a9fb97d6 · outbound

This paper cites Uniter: Universal image-text representation learning.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Uniter: Universal image-text representation learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.879541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:f63e8239dd38962af1d69f06f2c6a3e7e335964bd3a8ad0ae99ded3244e7d32f

Observation 0176d87a-d831-4b34-8df7-7bed01c48fb7 · outbound

This paper cites Towards natural language-guided drones: Geotext-1652 benchmark with spatial relation matching.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Towards natural language-guided drones: Geotext-1652 benchmark with spatial relation matching

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.892489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:ac364be4fbee4f84fd67956bf25b4eb5cb5350a793b7a68ad8fbae67f4c4faab

Observation 72bca8c7-2915-4eac-99de-947727dee97b · outbound

This paper cites Sam- ple4geo: Hard negative sampling for cross-view geo- localisation.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Sam- ple4geo: Hard negative sampling for cross-view geo- localisation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.886177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:7cd1f2815d464ce1a98742da905adb02f0ea70d75d2ae0ad8f504c76b3b651e7

Observation aebc0807-66f0-43f9-a787-bc8470014046 · outbound

This paper cites An empirical study of training end-to-end vision-and-language transformers.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization An empirical study of training end-to-end vision-and-language transformers

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.858426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:221c7c6ec6b3d99f2d9bfbe5e071ff03e1b64e5a399b1f69f1a6c9f57c63be58

Observation a1f6b41e-f0bc-4413-9e21-d19700714837 · outbound

This paper cites SimCSE: Simple Contrastive Learning of Sentence Embeddings.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization SimCSE: Simple Contrastive Learning of Sentence Embeddings

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:48:23.738089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:64100d966fa4e68427808e93a27f6c18c9d7f2d6e56049c1c869c3b6868efddf

Observation fbba1b2c-278b-4e3a-aaf6-1041f7aafe93 · outbound

This paper cites Skysense: A multi-modal remote sens- ing foundation model towards universal interpretation for earth observation imagery.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Skysense: A multi-modal remote sens- ing foundation model towards universal interpretation for earth observation imagery

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.861349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:bf00ff4485b834fb654948aa9278f0dc2d6d9259a759c364bfed8a0446ab1236

Observation ea5c2677-31b9-460b-abf0-78f7a044a282 · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Lora: Low-rank adaptation of large language models.ICLR, 1(2):3

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.864215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:8d83e7dd816c7df3c914b7d0a99e1718eb1edb3db4dc2d18ac4dd3f4455c65ac

Observation 806b1ad1-5c8b-4d0a-b1a4-90c8271e8821 · outbound

This paper cites Cvm-net: Cross-view matching network for image- based ground-to-aerial geo-localization.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Cvm-net: Cross-view matching network for image- based ground-to-aerial geo-localization

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.855540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:e2d576b9a3eae8e79ac3ce9fb4e41f97225995a8661057937fb19e0ab3ebc424

Observation 1b7bcdb3-3508-474e-b609-30cbbcfd138f · outbound

This paper cites Llm2clip: Powerful language model unlock richer visual representation.arXiv preprint arXiv:2411.04997.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Llm2clip: Powerful language model unlock richer visual representation.arXiv preprint arXiv:2411.04997

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:03.142149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:a2f8a864d359620f4f26ceafe8cb30fe2a49d17f9ff26bb8c25869a8fe8c637b

Observation 30baefd4-4190-4a6a-ac33-3b71ddb4b5fc · outbound

This paper cites Perceiver io: A general architecture for structured inputs & outputs.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Perceiver io: A general architecture for structured inputs & outputs

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.850110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:eb017846a5f731663e7e4af58f16887983ec2a43a770062e562f90fb75a570d3

Observation 1e1a664b-3957-4cea-9d5a-f04e26296ff4 · outbound

This paper cites Vlm2vec: Training vision-language models for massive multimodal embedding tasks.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Vlm2vec: Training vision-language models for massive multimodal embedding tasks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.794618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:c991e66ed4b81b15bc55c4b41e498526a580f949f532a92e46efcf9c03bf91ff

Observation 4f257d07-96d8-4847-b397-1dba4e768e66 · outbound

This paper cites Adaptive latent diffusion model for 3d medical image to image translation: Multi- modal magnetic resonance imaging study.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Adaptive latent diffusion model for 3d medical image to image translation: Multi- modal magnetic resonance imaging study

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.791415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:85a795beaa0e89749233b6ef7655693ae2892e49da947b03b858bb3824f6439b

Observation cf5e291e-01f4-4f18-9d1f-1b56f7e49ca1 · outbound

This paper cites Vilt: Vision- and-language transformer without convolution or region su- pervision.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Vilt: Vision- and-language transformer without convolution or region su- pervision

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.844184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:c93c88be5d62a4be0a7811c4c914bc25401e4ec25f6c1d5f76d1e87c71556aea

Observation f86c3c47-9acb-4dae-b770-2f764405f7b6 · outbound

This paper cites Multi- modal data-efficient 3d scene understanding for autonomous driving.IEEE Transactions on Pattern Analysis and Machine Intelligence.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Multi- modal data-efficient 3d scene understanding for autonomous driving.IEEE Transactions on Pattern Analysis and Machine Intelligence

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.847140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:c3f4be3e82824d6ecb6b9471725056a2ae705fb5f8b3898076968ccdb08cf1b8

Observation 6fe628fc-1c29-4371-8f1b-992d24d40711 · outbound

This paper cites Align before fuse: Vision and language representation learn- ing with momentum distillation.Advances in neural infor- mation processing systems, 34:9694–9705.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Align before fuse: Vision and language representation learn- ing with momentum distillation.Advances in neural infor- mation processing systems, 34:9694–9705

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.767551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:f595668232002a63fee9eee8fe942d5c3c481f34d8fe66de8bd5a4bb65410e08

Observation dcaf2212-d393-4b30-8575-59120121086c · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.867668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:d592a60227ec1d6feeec491defeab1c7e01094d2cae442067ce74c85b88f07d1

Observation 7f773f39-253c-463b-9752-157b2db5f67a · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.835238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:facf08add7bc4c16276c7b9c11fd99d5ae56ae9ad76ce0609ee1f53cc353d9b3

Observation 13ac5c34-8a8f-4d67-b03e-e2dceb445798 · outbound

This paper cites Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:06:03.125221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:6a1b9a785ccbb855ef45c086bb698016dd00fadf0be8e7f48c76e954ab02c1ca

Observation 660e612c-824f-440c-b605-90f21c299fd4 · outbound

This paper cites Cross-view image geolocalization.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Cross-view image geolocalization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.838274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:abaf02889b01dd737ad9c5ac2325315d5ed441accab41525cd5bf11cf7b2e757

Observation 4de0be30-099f-46e8-8ef3-cc78178eff8d · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.825426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:9914ceb0125d87c0e7bbc9680b0a6638a16033b025972e899af5589f1ac31ed4

Observation 547df760-e92e-4d22-aadb-c9ee28aea54e · outbound

This paper cites Lending orientation to neural networks for cross-view geo-localization.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Lending orientation to neural networks for cross-view geo-localization

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.907518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:e0b9062c5aac5fe809cece3c6995aa1956e076f4eefa1ddea00d60edc667a264

Observation f97c42fd-d0ac-442f-9219-061b4b0f42bc · outbound

This paper cites Delving into multi-modal multi-task foun- dation models for road scene understanding: From learning paradigm perspectives.IEEE Transactions on Intelligent Ve- hicles.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Delving into multi-modal multi-task foun- dation models for road scene understanding: From learning paradigm perspectives.IEEE Transactions on Intelligent Ve- hicles

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.822243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:d93bc2f19db5372739bacce63e476480f3719ab2c6da7a19ced36bdd0281e9e8

Observation bf5d5300-f2f7-4088-a80f-820020dbe578 · outbound

This paper cites SmolVLM: Redefining small and efficient multimodal models.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization SmolVLM: Redefining small and efficient multimodal models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:23:51.804797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:5351fb0a54201ab78081fef2631b1f64ea6fe5f58e1049c40baae3afe4a7ceb4

Observation 7ab15c9b-991f-40d3-8155-23cb3dbd2f4f · outbound

This paper cites In defense of dual-encoders for neural ranking.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization In defense of dual-encoders for neural ranking

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.828436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:f93b753f60ac774a6313e86bd24b019d59294ad19f1909aaef3fbbc312d5c082

Observation 198a5989-459b-4fe9-93e9-f6a7de73db78 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Representation Learning with Contrastive Predictive Coding

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:06:03.094262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:b112201ae1019d8d213f3f5053558390adb52a988c4d30b9027c37a9d107f91c

Observation 1bef32d1-59fb-4554-bae2-964262726431 · outbound

This paper cites Training language models to follow instructions with human feedback.Ad- vances in neural information processing systems, 35:27730– 27744.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Training language models to follow instructions with human feedback.Ad- vances in neural information processing systems, 35:27730– 27744

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.818883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:5d679413c796e598ae5347e26310304fc1b89a7433d578887fbcad91710bc1a9

Observation 68229d8d-e24b-4513-805c-b9d78f5dc78e · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Learning transferable visual models from natural language supervi- sion

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.921666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:49626268d222eae86884f75177ea1c893b546e2658e501f7077591df31280992

Observation a43f8fce-ba9c-4497-808f-b41ee9b78c9c · outbound

This paper cites Cross-view image synthesis using conditional gans.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Cross-view image synthesis using conditional gans

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.812052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:966b937fb329aca27655b28e08be8d0381449cf79c47ddfe92f8fd2d2ce076e8

Observation 6cd44da2-16a8-4616-99a0-aabbb97e62bc · outbound

This paper cites Multi- modal vision pre-training for medical image analysis.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Multi- modal vision pre-training for medical image analysis

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.815525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:29fc0ad498c7f61d6480c1c5af2d7cc14a02bf5c17685ab9dfbcf48b5a86c744

Observation eea14c1a-80f8-4553-b7e5-5a8da5896006 · outbound

This paper cites Toolformer: Lan- guage models can teach themselves to use tools.Advances in Neural Information Processing Systems, 36:68539–68551.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Toolformer: Lan- guage models can teach themselves to use tools.Advances in Neural Information Processing Systems, 36:68539–68551

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.831959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:db3f764c82d5d6c7e8c266d04e011f56adedb283f863d4ce66e30fe859edfcc3

Observation 2015a3b4-c445-4e5e-9c86-15ea364112d4 · outbound

This paper cites Spatial- aware feature aggregation for image based cross-view geo- localization.Advances in Neural Information Processing Systems, 32.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Spatial- aware feature aggregation for image based cross-view geo- localization.Advances in Neural Information Processing Systems, 32

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.841392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:fd8b3a221eae6c5d90aa15ca21159e89e6350ef6e29ea9e215f541ea5039546f

Observation 164ae84f-a7da-4211-ad61-f88ae0659f26 · outbound

This paper cites Optimal feature transport for cross-view image geo- localization.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Optimal feature transport for cross-view image geo- localization

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.852865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:a7bf04f71aceff5c03fc443805748535706a1f367011c993b12d683c6b6f3eb5

Observation a98338c8-6ea8-4977-88cf-361a9dc2b3e2 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Gemini: A Family of Highly Capable Multimodal Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:06:03.087498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:0ec8f1f5b6915a672c6290b3828b876036be8277ec6723af71009d77d86dab9e

Observation c33741d6-c638-41c4-9229-84ac3a8dec2a · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Attention is all you need.Advances in neural information processing systems, 30

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.889461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:50af44f74889011e1073fcd35ae4b230c03d7c43404042b51a4fc40eeea235a9

Observation 42b1840e-9d3c-4386-b0e7-af67e2b51ee1 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:06:03.068874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:7b48bfee0bed5904531e4ab63e602ac1bd5e84a50edec6a5808f068533f13d60

Observation 1fc1c867-a1dd-4fab-a466-0cb8a3bb4141 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:06:03.102161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:0345aff11e965a1ef06589d6087d903d0338921bf553222bb5333798a4f3615d

Observation 3e5070c8-a9d0-4c2d-9062-e96810655f7e · outbound

This paper cites Image and object geo-localization.International Journal of Computer Vision, 132(4):1350–1392.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Image and object geo-localization.International Journal of Computer Vision, 132(4):1350–1392

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.806187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:39c2be7e3e54728bd92f90f02f73e29eb6154b999a10141c246de60e1189d13b

Observation aaebdf15-94e8-42dd-80f4-0b2878d9ee40 · outbound

This paper cites Wide-area image geolocalization with aerial reference im- agery.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Wide-area image geolocalization with aerial reference im- agery

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.918202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:b4fbd1bdd9066ab3ec0903742cc2c55b1e70111b999284b158b7058ae7700da8

Observation 62c6e12e-80e3-436c-a22b-0f5bc5a0d520 · outbound

This paper cites A semantic-enhanced multi-modal remote sens- ing foundation model for earth observation.Nature Machine Intelligence, pages 1–15.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization A semantic-enhanced multi-modal remote sens- ing foundation model for earth observation.Nature Machine Intelligence, pages 1–15

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.911101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:b1bca2fe63fcf82b239bfd8631878913ecf916f45086a71ecdf8d4e121dca4d3

Observation 509df574-1fa4-4c7a-88dc-1dfbe99033de · outbound

This paper cites Cross-view panorama image synthesis.IEEE Transactions on Multimedia, 25: 3546–3559.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Cross-view panorama image synthesis.IEEE Transactions on Multimedia, 25: 3546–3559

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.809078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:8dada286c373447ca1c9f1e62f32f833d3335ae3fc0846aa7a64f9e2b4960eed

Observation ec7bd879-6696-495f-b311-be91c2c1bb44 · outbound

This paper cites Cross-view geo-localization with layer-to-layer transformer.Advances in Neural Information Processing Systems, 34:29009–29020.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Cross-view geo-localization with layer-to-layer transformer.Advances in Neural Information Processing Systems, 34:29009–29020

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.797761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:32d1b6d655162f190701bb36a33490d51eafa3fc9492444ee979d435641eaf33

Observation 0f65dcb1-f372-4242-9cb2-d7a1dd9ba1ce · outbound

This paper cites Cross-view image geo- localization with panorama-bev co-retrieval network.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Cross-view image geo- localization with panorama-bev co-retrieval network

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.800578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:0da093636bbf1038de3aacecb7b317c3a0d86999da242b0472014820ceb1012f

Observation b6da84ac-ff7d-4c7f-b16c-f61498751b53 · outbound

This paper cites Where am i? cross-view geo-localization with natural language descrip- tions.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Where am i? cross-view geo-localization with natural language descrip- tions

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.803510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:76f82fed8cb85ee3d3212097efa90d8243176e3c155669fdab0aa315aeab2add

Observation c10329b8-9170-4f21-bdbf-22f4bf68703b · outbound

This paper cites Multi-grained vi- sion language pre-training: Aligning texts with visual con- cepts.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Multi-grained vi- sion language pre-training: Aligning texts with visual con- cepts

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.914540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:af856909b3a7ae4c7843242631bb5dd0a7f0bb66ac3621fe41d508181dab2107

Observation e86013c2-133b-4adb-83bd-e2e6e0c52c70 · outbound

This paper cites Sigmoid loss for language image pre-training.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Sigmoid loss for language image pre-training

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.785094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:38ba3e7777d3d21a57cf53daeaf4687332f50ed7b3a00fb9dab3f21ce1cce4bb

Observation 0cadc3b4-1c7a-4a20-adfb-b0b810c48d0b · outbound

This paper cites Cross-view geo-localization via learning disentangled geometric layout correspondence.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Cross-view geo-localization via learning disentangled geometric layout correspondence

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.788286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:eb792be9608cd4edae74f9f74556a3e3e7260dbdae25627d98b3dbddf37c6273

Observation 3e971db1-650e-4db2-a25e-5bac26180ba5 · outbound

This paper cites Cross- view image sequence geo-localization.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Cross- view image sequence geo-localization

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.778170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:277230a3988ccc1d58e8a3130a98973021c8a003ad1d0c551690590e5f25f507

Observation 1033dea7-729c-45b1-bbe6-97fec9b30bb5 · outbound

This paper cites Geodtr+: Toward generic cross-view ge- olocalization via geometric disentanglement.IEEE Trans- actions on Pattern Analysis and Machine Intelligence.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Geodtr+: Toward generic cross-view ge- olocalization via geometric disentanglement.IEEE Trans- actions on Pattern Analysis and Machine Intelligence

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.904271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:bc913359fc6eb27c23ac29a55af2a2cb730f6044870274fef20ea6b971afce53

Observation 3980a872-6fa9-4135-bc5b-e0048f1f15ed · outbound

This paper cites Vici: Vlm-instructed cross-view image-localisation.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Vici: Vlm-instructed cross-view image-localisation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.774946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:efc308175150afb33a5e7a6eeb958bfca41d744174ef2c95cfc7b5410d236db3

Observation 3e47ac36-6c49-4238-812d-c918c76cddb1 · outbound

This paper cites University- 1652: A multi-view multi-source benchmark for drone- based geo-localization.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization University- 1652: A multi-view multi-source benchmark for drone- based geo-localization

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.781740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:1c92f0a49ae72df6c7c67843b893334fd960d75b94c4fed01168350dd695bb25

Observation 5aa9a884-8122-4486-b5e2-d5420d0a2125 · outbound

This paper cites Vigor: Cross- view image geo-localization beyond one-to-one retrieval.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Vigor: Cross- view image geo-localization beyond one-to-one retrieval

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.896304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:5d816096af7c09b791ec6128ecee74f9ab67e9d07376c87e1b41801be61bbc40

Observation dc273044-65d1-408a-b528-d38527474df2 · outbound

This paper cites Transgeo: Trans- former is all you need for cross-view image geo-localization.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Transgeo: Trans- former is all you need for cross-view image geo-localization

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:30:10.771439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:16d93dcbaead3aad41d47d49d62a99eaff764839966914b3a51f7943cf70e607

Observation aff193ed-4a57-47eb-88ca-aae43219a3c2 · outbound

This paper cites Simple, Effective and General: A New Backbone for Cross-view Image Geo-localization.

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization Simple, Effective and General: A New Backbone for Cross-view Image Geo-localization

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:03.110441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:39:17.229872Z digest=sha256:f12495cb4368cb16570fa17168514ca41a7ec9fa70fcefa9953070b662599c77

Pith citing papers

No inbound Pith citation observations are available.