Pith. sign in

Paper Citation Record · LEDGER

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models

As of 7 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 6 inbound Pith citation observations for arXiv:2506.01277.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01277 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:52:10.218789Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:12:10.857306Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T23:06:20.239381Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact4
  • verified fuzzy20
  • unresolved14
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5b870ec5-2ecf-4f1f-9453-a203fb34e9ec · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Learning Transferable Visual Models From Natural Language Supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:08.326032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:08.326032Z digest=sha256:254106d8eab177fef676c0ec30c5fded59e54a97cd938060d31b4411cc8d8192

Observation 6a14bccb-dcf5-45f6-b2a3-7a98fa9acb10 · outbound

This paper cites an unresolved cited work.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:52:13.349011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:08.362373Z digest=sha256:3e3b08eed007bb69327d3e2e86516782c0b6fe0afa0df8afd3e94dfc44e6cfd2

Observation 1141b9f7-2500-4a76-8e86-177824a35b62 · outbound

This paper cites PlaNet - photo geolocation with con- volutional neural networks.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models PlaNet - photo geolocation with con- volutional neural networks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:08.387126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:08.387126Z digest=sha256:4d7163446f691224d77a7fe3c255a31211dea2262f3dd75a1c6751ae58db2a61

Observation 148ace07-28fb-49c5-85a7-6ba3ad394ec2 · outbound

This paper cites PIGEON: Predicting Image Geolocations.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models PIGEON: Predicting Image Geolocations

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:52:10.869770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:08.433343Z digest=sha256:347f049502a2e99f683d26ad0f6770398a99583c93c34fd51cd0ca0147330c1f

Observation fee7286f-50b1-4b48-a43b-13fa9eb5c07f · outbound

This paper cites Image-Based Geolocation Using Large Vision-Language Models.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Image-Based Geolocation Using Large Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:08.551981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:08.551981Z digest=sha256:8a1770b830fda104e9285986d64cfbd33364107a5807c813fd2447ec321aa97c

Observation a2dc9744-dbad-4f52-852a-7bfcb6c5134f · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:08.601226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:08.601226Z digest=sha256:2206fd806bf53f4ca686f9510df04640b6fa06177d71707b1fcd6e8eeda35fba

Observation 7b13da83-2661-40bd-b0f2-60b00ee1c4f7 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:08.630800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:08.630800Z digest=sha256:fda354474d5ba3702c8ca0a0b02ee5c5f30efe73f29473edb6afe52afa840843

Observation 30c03c75-74da-45cd-a91b-d1218652f3af · outbound

This paper cites Neural network ensembles.IEEE Transactions on Pattern Analysis and Machine Intelligence, 12(10):993–1001, 1990.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Neural network ensembles.IEEE Transactions on Pattern Analysis and Machine Intelligence, 12(10):993–1001, 1990

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:13.244819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:08.677012Z digest=sha256:afc609672505c4eca7966db2bc3bb8e5ac8d294ec1df302f36eb2e1350231c40

Observation a64f0ec5-2882-4c39-8759-e00d3c24ef5a · outbound

This paper cites Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:08.724712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:08.724712Z digest=sha256:e1f5d09f4eb44b83de75dff2b11ca7c654292cbee0a92f1746e56b1a3ed8ae73

Observation 8dcf04f6-b1eb-42bb-9bd8-6608e1952def · outbound

This paper cites Where We Are and What We're Looking At: Query Based Worldwide Image Geo-localization Using Hierarchies and Scenes.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Where We Are and What We're Looking At: Query Based Worldwide Image Geo-localization Using Hierarchies and Scenes

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:52:10.616398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:08.764186Z digest=sha256:701b3ef19b96b45b97cc7715733300d22a2f3179eec4210edec7db946cbca65b

Observation eb76507a-c259-486b-a3ce-9e3b118146cc · outbound

This paper cites The Mapil- lary Vistas dataset for semantic understanding of street scenes.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models The Mapil- lary Vistas dataset for semantic understanding of street scenes

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:13.100674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:08.801352Z digest=sha256:5bb856225740a8fd35a46a68420a8141a83c7bc1db76306098e4f62fc5471f76

Observation 25452abf-1357-45ad-88ff-b01d1e19727f · outbound

This paper cites GeoNames geographical database.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models GeoNames geographical database

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.947008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:08.826571Z digest=sha256:f7184bac5325ba734ae28764f4537600bd8dfd47b18f1bee45286870dca13402

Observation 5169bda4-4f7d-42c3-a9a2-4dcc23558192 · outbound

This paper cites Qwen2.5-VL Technical Report.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Qwen2.5-VL Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:09.047711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:09.047711Z digest=sha256:7b5310c25a166536d53c97cf567b4f96796e5bc12936b0782076a2543b23f273

Observation 54ddb3a1-3421-4792-a162-fd35c6d2a87d · outbound

This paper cites Decoupled Weight Decay Regularization.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Decoupled Weight Decay Regularization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:09.081528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:09.081528Z digest=sha256:60b7c6924b7d2f2d52a80661b0503c2a72b43be6ea00e2c61d179fda216d5e72

Observation 9ad4f2a6-8671-4596-87c9-e9dc77b2b336 · outbound

This paper cites OpenStreetView-5M: The Many Roads to Global Visual Geolocation.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models OpenStreetView-5M: The Many Roads to Global Visual Geolocation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:09.120531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:09.120531Z digest=sha256:f101f204a9adb79f233176abd9747435c401f7c5321f61bf53ef216c9f9b22f6

Observation 193b1e22-a3e6-49ce-8702-3dc16a604b34 · outbound

This paper cites The Claude 3 model family: Opus, Sonnet, Haiku.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models The Claude 3 model family: Opus, Sonnet, Haiku

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.868679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:09.193032Z digest=sha256:316fbcbad141ce2f453b0842b51e4f410b77f2988a0e09faeec35b43e929d51e

Observation 7631d502-cc44-4fdb-9cae-f977578e16e5 · outbound

This paper cites Visual Instruction Tuning.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Visual Instruction Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:09.261284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:09.261284Z digest=sha256:5a75e3b5f042feaae7ef47a4e39b89264d0d2be8d587fa9479c092090ed94786

Observation b8c927a4-7674-4873-acef-447b02dfb975 · outbound

This paper cites Mistral-Small-3.1-24B-Instruct-2503.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Mistral-Small-3.1-24B-Instruct-2503

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.768597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:09.302998Z digest=sha256:6deeb6a2fe0ed16c6fb62c6c8fc2789ee6425d08651b6008916a294531d08b94

Observation 00ff3373-43a2-4640-9a60-16d45423daa3 · outbound

This paper cites Revisiting IM2GPS in the Deep Learning Era.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Revisiting IM2GPS in the Deep Learning Era

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:09.371329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:09.371329Z digest=sha256:36cf39cd33d56c5f438452458ec25d2ca6ea0c9e9fe0dbaf430a9439afd13fad

Observation fe163f24-9bf9-4a1e-9068-23c0b9512cd1 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:09.515725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:09.515725Z digest=sha256:16a7d17d980fb7c5a8a4bbb0f6b70c15381f2805105b61bc9402ce42ca3193b7

Observation 4c703f5e-d585-4cf8-9785-e4d83a4ec20f · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 24

Resolution
malformed identifier
no resolver link, observed 2026-08-07T11:52:09.555930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:09.555930Z digest=sha256:2822b029fb89e667fc3df84c931681b5c7a729ea02917e150196c7fc711b342d

Observation e70fda6e-38f2-403b-a5b8-43f8a2764ae7 · outbound

This paper cites Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:52:10.415712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:09.473416Z digest=sha256:2b7808eee71f521b52cb260b75d190bfbc4e4e4612eaf0d6a8d96e9ace4962b8

Observation 7dbbf803-2ed1-4664-ac0f-00481042e4bb · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.675504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:09.594635Z digest=sha256:4c53572bda21d0fe23666480cd3ca998aa0c215bcc0fd97aefa580282c44084d

Observation 947e2810-6378-40e2-b3cf-419d7b1409a3 · outbound

This paper cites Guidelines: • The answer NA means that the paper has no limitation while the answer No means that the paper has limitations, but those are not discussed in the paper.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the paper has no limitation while the answer No means that the paper has limitations, but those are not discussed in the paper

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.601524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:09.627183Z digest=sha256:3f3a32cde0c24590ed3ff377c67017f7e3d0a4cd4a83a3a71584c24cb487e345

Observation 9bcd7b9a-f880-435d-aec0-4ad3a7a11462 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.486517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:09.687952Z digest=sha256:49c9e9bf3383b6c176f48a775619f6c0433b1d0d795cb8e35181095b3f224b74

Observation 62811774-990e-4f8c-9518-4fa0e1845b78 · outbound

This paper cites The appendices include hyperparameters for SFT and information about computational resources used.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models The appendices include hyperparameters for SFT and information about computational resources used

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.393269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:09.751346Z digest=sha256:7d7116f032a13aac0676a1bd0817df43b448c005bba7107c224d83895e2139d3

Observation 7dc63ab1-963b-4bcd-b217-e92994f77b42 · outbound

This paper cites Our appendices provide detailed instructions regarding implementation, hyperparameters, and experimental setup to facilitate reproduction.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Our appendices provide detailed instructions regarding implementation, hyperparameters, and experimental setup to facilitate reproduction

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.301578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:09.792357Z digest=sha256:9f7452603426b669c56e0e4b5d7bc99da2d109f8914a187565d313b82f289366

Observation 1dba0e3d-4065-44e2-a2ff-01d772224858 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the paper does not include experiments

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.220821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:09.842854Z digest=sha256:05d92f64013490c51ff8cc5b622e67d937fd8648a3408bf4797c53953eea4c79

Observation 5cc46a9c-3c83-45bd-a47d-a8079eb9fd12 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the paper does not include experiments

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:12.110456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:09.878555Z digest=sha256:55209983230accb7f4794a32bd656056e92e3a613952160a4ee9b9ff13e42313

Observation 180e1af7-572b-4d03-8655-1beb3d7b8761 · outbound

This paper cites We also report model sizes and memory requirements.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models We also report model sizes and memory requirements

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:11.988746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:09.888164Z digest=sha256:1605e9c3db5400e1a8772cd50cbccef9edadabd6d0e70b75f289decb47fa358b

Observation 13c45ece-fa9c-4e7a-836b-b714c48619b8 · outbound

This paper cites We use publicly available datasets, acknowledge relevant prior work, and are transparent about our methodologies.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models We use publicly available datasets, acknowledge relevant prior work, and are transparent about our methodologies

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:11.886563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:09.925377Z digest=sha256:4b39398367304b1bf714032e9607fcb93258377a53125fd1cb5f494889a369ad

Observation 54276147-772a-4883-9c90-95322fe4ce64 · outbound

This paper cites Guidelines: • The answer NA means that there is no societal impact of the work performed.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that there is no societal impact of the work performed

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:11.755852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:09.978969Z digest=sha256:93eec4e3f312f98a043153d3de0129d90e8409863bd934818ddee2c662566f55

Observation bbf48e5d-e190-40dd-9d6b-1ed9331aa26a · outbound

This paper cites 27 Guidelines: • The answer NA means that the paper poses no such risks.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models 27 Guidelines: • The answer NA means that the paper poses no such risks

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:11.621881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:10.027853Z digest=sha256:0f961690eb4970ca85f31229b89fa18f8d0389c419a2e7393016ccf03b3ae72d

Observation 3ef73617-d963-4be5-98a5-33ed5ad543e8 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not use existing assets.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the paper does not use existing assets

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:11.507759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:10.068261Z digest=sha256:80f317c264b54c7433877d043655fb7ab903bdb860f1a72a2cb5d928d149186c

Observation 8125e676-e808-4062-8bd4-3cebc78c0535 · outbound

This paper cites This documentation will be released alongside the dataset.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models This documentation will be released alongside the dataset

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:11.310523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:10.112687Z digest=sha256:3d6024eb18ba15eabf23119d54c6f982ea85f01d93869c4dad653c7ad9252248

Observation 04d645b3-ad62-46d9-a66e-d21e141be845 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:11.158826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:10.143213Z digest=sha256:79b6bb36de50a82db1ba2cc8a5c754c64b29409bcf9e0ffb5e97ce7ad969aa13

Observation c14f16c0-6101-481f-82fd-6053b2217ae5 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:10.997861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:10.218789Z digest=sha256:8848d9d1cdc8815c691684eb25727a291ca808977165fd657e2dd72e4e61c37f

Observation 2183609d-e1a8-486a-821c-40d15e3a9070 · outbound

This paper cites G3: An Effective and Adaptive Framework for Worldwide Geolocalization Using Large Multi-Modality Models.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models G3: An Effective and Adaptive Framework for Worldwide Geolocalization Using Large Multi-Modality Models

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:52:10.788245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:52:08.501464Z digest=sha256:d1a4003a5f57776ebe9812d2d9288928c6cf0a5109b4790ab35a389da3c26e44

Observation b42b2fe3-b71a-4bb4-a522-2881740e889e · outbound

This paper cites Gemma 3 Technical Report.

GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models Gemma 3 Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:08.967423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:08.967423Z digest=sha256:f213e41e41316f719bf4ae360de3c913d0e5a375fd1ace0a836abbadf6afb499

Pith citing papers

Observation 5da9a71f-b019-499b-aeee-179f85ab3b52 · inbound

A Global-Local Cross-Attention Network for Ultra-high Resolution Remote Sensing Image Semantic Segmentation cites this paper.

A Global-Local Cross-Attention Network for Ultra-high Resolution Remote Sensing Image Semantic Segmentation GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:10.857306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:12:10.857306Z digest=sha256:13ed3664ea89375d821721724975207337f9f931cd6835b26cdf8032f7e4fc8b

Observation c1b68c6f-60f2-4d40-b263-be1b4624e2e8 · inbound

Building Lightweight Semantic Segmentation Models for Aerial Images Using Dual Relation Distillation cites this paper.

Building Lightweight Semantic Segmentation Models for Aerial Images Using Dual Relation Distillation GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T23:00:29.112400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:00:29.112400Z digest=sha256:1c8b453a0c6b0553b90903c745459a43ae2531025b5850f37085934b1110c377

Observation 0a34a1ef-010f-4b43-8de7-cf7d9ee70961 · inbound

GDGS: 3D Gaussian Splatting Via Geometry-Guided Initialization And Dynamic Density Control cites this paper.

GDGS: 3D Gaussian Splatting Via Geometry-Guided Initialization And Dynamic Density Control GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:24:31.913301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:24:31.913301Z digest=sha256:0bf055d688a11e9a2f94bd3eab57c5fcfccc354d42253a813a16283ac3768e44

Observation 4aab417b-4cd4-4650-9ded-b40c41c22580 · inbound

From Pixels to Places: A Systematic Benchmark for Evaluating Image Geolocalization Ability in Large Language Models cites this paper.

From Pixels to Places: A Systematic Benchmark for Evaluating Image Geolocalization Ability in Large Language Models GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:50:47.271597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T23:49:59.795575Z digest=sha256:6c741c0dec4da10197273fbe4fe5fb16b91de0464fb84df72947ff16ddbae1bb

Observation 556d334e-529f-4853-a0ff-ca369f038b59 · inbound

Do VLMs See What Sensors Feel? A Scalable Expert-Guided Design for Wheelchair Accessibility Assessment from Street View cites this paper.

Do VLMs See What Sensors Feel? A Scalable Expert-Guided Design for Wheelchair Accessibility Assessment from Street View GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:20.241225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T14:42:47.815192Z digest=sha256:8d88ee92440dd9cef369698873d3058f83e38e04b19e51953806f68e5c84da7f

Observation 33241b1a-4fed-42af-80bf-13ec998b5cfb · inbound

DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization cites this paper.

DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-30T22:45:19.083071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T22:45:19.083071Z digest=sha256:7d939089483238f12f665c2da1e2d1f789eebc23e2af86ac34cfddad8b9b87dd