Pith. sign in

Paper Citation Record · LEDGER

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale

As of 5 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 0 inbound Pith citation observations for arXiv:2604.12159.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.12159 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:37:16.031502Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

79 of 79 outbound references displayed

  • verified exact14
  • verified fuzzy57
  • unresolved6
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9821f85b-dd29-44af-85bf-fef6d3b38c67 · outbound

This paper cites Openstreetview-5m: The many roads to global visual geolocation.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Openstreetview-5m: The many roads to global visual geolocation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.234244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:079ff713b7edbeeba4595c038e1d8391c1c39ac44ac906d64bcaacf6014a105c

Observation 1a55a4e6-c8e3-402a-9ae2-9ec7c4c77ed4 · outbound

This paper cites Qwen2.5-VL Technical Report.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Qwen2.5-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:11:02.632798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:52e1f66b8b8cc3e6a3e4d081692dc8bb61e6c9af086d225a33f8e9223c060c08

Observation 3fd33036-f77b-49b1-a299-20a85cc9c841 · outbound

This paper cites Is space-time attention all you need for video understanding? InICML, page 4.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Is space-time attention all you need for video understanding? InICML, page 4

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.045260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:2b31cbd903993b2ae4516bdd18517db181a4af1cf4a11c28773bac3a646e42e3

Observation da7f6acc-bf63-4078-ae4b-7982894953ae · outbound

This paper cites GAEA: A Geolocation Aware Conversational Assistant.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale GAEA: A Geolocation Aware Conversational Assistant

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:02.494381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:905c2abc7ecec2427388378de59628a5e23e29afd6e8fca68a064e984df099ba

Observation c2122a78-d0bb-40f5-b11a-77cb588ff1e0 · outbound

This paper cites Where we are and what we’re looking at: Query based worldwide image geo-localization using hierarchies and scenes.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Where we are and what we’re looking at: Query based worldwide image geo-localization using hierarchies and scenes

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:09.994761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:8e12632df5ef87a0fac840e8871c972282a2c7817b314292f881bbbf01fab2a1

Observation ffc09fbc-6f57-4d6e-a140-6e90acb196b1 · outbound

This paper cites Sam- ple4geo: Hard negative sampling for cross-view geo- localisation.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Sam- ple4geo: Hard negative sampling for cross-view geo- localisation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.135660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:e044a6898eea8ca8b5a973142c8c159dc21cc9bdc41d218547dd7e5f3044e7b4

Observation c036f975-cafb-4a28-9499-7f3beb61154a · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:11:02.603817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:48bd3a5ebb69ec6670d576cfc9b08fefaae386330f4d08a6e05d19a8c08a84a1

Observation be4df011-292c-41fb-aa99-18b756724f86 · outbound

This paper cites Computing discrete fr´echet distance.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Computing discrete fr´echet distance

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.172233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:7c007feaad051fe72da7846cdb98b28009877c11291531730a2d7fb885fb55e4

Observation d23cbc21-faab-445a-9afa-89c86d3e583e · outbound

This paper cites Condition-Invariant Multi-View Place Recognition.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Condition-Invariant Multi-View Place Recognition

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T23:25:15.858185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:f89bffb8fd55d8af59c6b76914261a8f08fe0743bdf94d6811503443eb24e8bc

Observation 55b44c1c-fd3e-4e0e-8aec-8d7030a6bf22 · outbound

This paper cites Xi-net: Transformer based seismic waveform reconstructor.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Xi-net: Transformer based seismic waveform reconstructor

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.120313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:1ca5e2c440c649247b2a97d41a82cbaaa35273333b0b4d691c67539edc39c390

Observation 3c10c8e7-8bb4-4281-8f47-2b13ce4005f1 · outbound

This paper cites Garg and M.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Garg and M

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.125978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:61382b9ede3b2d03113b24ba28cf5ac319da46e73b579a776d16127a484f5306

Observation 56665802-6947-4fac-85f0-93985fa328d8 · outbound

This paper cites Learning Generalized Zero-Shot Learners for Open-Domain Image Geolocalization.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Learning Generalized Zero-Shot Learners for Open-Domain Image Geolocalization

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:02.576543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:e20118fa58cdd23cbbe444779c5f53adeff613df9bf0ae4d7e5a669379c25940

Observation 96ca759c-4f90-4cdb-99c3-3d27c88bd6c2 · outbound

This paper cites Pigeon: Predicting image geolocations.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Pigeon: Predicting image geolocations

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.176059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:735f066f58569d99f1033d78a34957a0e588695ef7dfd11cddb44c22bb79fbb8

Observation 9a479111-0507-4477-bef5-ea4e1bbd0590 · outbound

This paper cites Im2gps: estimating geo- graphic information from a single image.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Im2gps: estimating geo- graphic information from a single image

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.067130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:8657e459c0eb72b79b8dd5a490ebb9957d32a91505a8276691fe099ef677a1ba

Observation 03e5c462-7415-4343-99bc-6759f32571ca · outbound

This paper cites Deep residual learning for image recognition.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Deep residual learning for image recognition

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.083335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:cda96ed68038c46b9442a260ae5fc91839788122f49bb8519cfe31267c3c8d3a

Observation 55dc9d8d-b76a-48c8-8f65-269082db6759 · outbound

This paper cites Cvm-net: Cross-view matching network for image- based ground-to-aerial geo-localization.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Cvm-net: Cross-view matching network for image- based ground-to-aerial geo-localization

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.102957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:91eb39638d2bf090b1d104af7c0533e98d3f9dced111228818748b490bce5f34

Observation 5e5972db-d612-4f8f-b392-f5b6e8154560 · outbound

This paper cites 3d convolu- tional neural networks for human action recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(1):221–231.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale 3d convolu- tional neural networks for human action recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(1):221–231

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.142349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:212c61252ba21b377edbb9a7ac6ceb2b33fccdf2c8d73fe8310f8abb1e22693e

Observation 13bcc413-7249-4c48-a629-1ab9a34f42ff · outbound

This paper cites G3: an effective and adaptive framework for worldwide geolocalization using large multi- modality models.Advances in Neural Information Process- ing Systems, 37:53198–53221.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale G3: an effective and adaptive framework for worldwide geolocalization using large multi- modality models.Advances in Neural Information Process- ing Systems, 37:53198–53221

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.138969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:6e48f6afe73aab3f5212bcaee5d225a8a2941ec7ddd813978182e52feb6d01a3

Observation 6677e735-906a-49ca-92cd-410ee9e88da7 · outbound

This paper cites GeoRanker: Distance-Aware Ranking for Worldwide Image Geolocalization.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale GeoRanker: Distance-Aware Ranking for Worldwide Image Geolocalization

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-25T02:17:29.888408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:dce032155c62c76d0e6c35cce0bfa9d3333f48f9708e8169366a473614a198b6

Observation 163fc47b-a675-4f99-bd49-ea567918cd2b · outbound

This paper cites Adam: A Method for Stochastic Optimization.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Adam: A Method for Stochastic Optimization

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:11:02.460366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:e5cf703c941b5f9f0a7666470edcdcddd5d103908f803136abdb5e5eca62ad0c

Observation 3f1962f3-88d3-423f-9f18-eeb2e1b08ab9 · outbound

This paper cites Cityguessr: City-level video geo-localization on a global scale.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Cityguessr: City-level video geo-localization on a global scale

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.031229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:7dccc15e1826f587ce2de1b464b8ba462e82ee647d9991fd33c29a790f971739

Observation 7c511823-3dba-4130-9ad3-3511be810ce7 · outbound

This paper cites The benchmarking initiative for multimedia evaluation: Mediaeval 2016.IEEE MultiMedia, 24(1):93–96.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale The benchmarking initiative for multimedia evaluation: Mediaeval 2016.IEEE MultiMedia, 24(1):93–96

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.089553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:deacf6c5d467a6d74599594607721a47e287b2750377f6c90bbfb2e7d3cb7989

Observation 01274798-9a51-49c0-95f5-5d84a2a1168a · outbound

This paper cites Handwritten digit recognition with a back- propagation network.Advances in neural information pro- cessing systems, 2.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Handwritten digit recognition with a back- propagation network.Advances in neural information pro- cessing systems, 2

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.187334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:d15ceab282a87a8aa4b5c72be9f2b521f8308fba7f7c9c63ef6fc9598f99b1e1

Observation 87b2a349-d954-4524-b4b4-2af50db03fbe · outbound

This paper cites BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T00:14:58.269455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:9d851941035da7f062fc2fe8c33bfaa9a15b3bf1a62d9ccb5d85c5e57ee7b87c

Observation 4a8e8144-1f57-434e-8006-29e500f043c0 · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing systems, 33:9459–9474

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.168594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:155e2587d1042dab238e52c1cb38c9b836bd2621cc585a4ca01fcd46cc71ad93

Observation b03ac706-e594-452d-b4c5-3eb4b6ee1af7 · outbound

This paper cites Georea- soner: Geo-localization with reasoning in street views using a large vision-language model.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Georea- soner: Geo-localization with reasoning in street views using a large vision-language model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.231030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:8020c90e2db9db78aa8ee24041c92ff9f026f6c32fa1cf36bb9491ee146848e0

Observation 2acd3245-4478-447e-b102-5320ba74841c · outbound

This paper cites Lending orientation to neural networks for cross-view geo-localization.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Lending orientation to neural networks for cross-view geo-localization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.054742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:187c3241fe8444add4fbd4a32f1b7dd3c730b99abf84d7222c53e5de13769157

Observation 8897277d-ba04-4959-add9-006dd37008d8 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Swin transformer: Hierarchical vision transformer using shifted windows

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.079906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:ccd83c2cf7021c9d1ea99435b4afede35757726cd21b1bb9e3fc4ff9ad2c1863

Observation b6855385-8d9a-4a44-9768-45d2ad88efde · outbound

This paper cites Mereu et al.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Mereu et al

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.092885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:9b7498194f66de4539fe314915e3932a81a6e5d490aa57873b34137e43cb3f4b

Observation 861f9c45-f3e1-4a91-a776-29211e3b5701 · outbound

This paper cites ConGeo: Robust Cross-view Geo-localization across Ground View Variations.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale ConGeo: Robust Cross-view Geo-localization across Ground View Variations

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:02.482280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:f9cd7a00324b692c373c92c9b9a628be134bea69d6c8682507efc66788607927

Observation 37c45461-28c3-4851-a987-984b78ff0629 · outbound

This paper cites Mish: A Self Regularized Non-Monotonic Activation Function.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Mish: A Self Regularized Non-Monotonic Activation Function

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:02.509102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:493281d48eff5d16b677618f0f9f4922bde3b0954e342e329f44a4718a70a6da

Observation f34e7bd2-823e-40ec-82d0-0616c8c4bc54 · outbound

This paper cites Geolocation estimation of photos using a hierarchical model and scene classification.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Geolocation estimation of photos using a hierarchical model and scene classification

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.183642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:b33bd5faf347e5a9a8a61fa724178070d209244ac26fcd7d5512897494bb1384

Observation 24b1f072-89c6-4f16-8633-b5359be3dae7 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale DINOv2: Learning Robust Visual Features without Supervision

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:11:02.533673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:c16fa72ca18b96a78b7deaab8b61b0bb4b7dd2711481f1c8086bfa07a9884779

Observation b3b85a14-6bae-47a5-bf93-99ce3807da00 · outbound

This paper cites Pytorch: An im- perative style, high-performance deep learning library.Ad- vances in neural information processing systems, 32.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Pytorch: An im- perative style, high-performance deep learning library.Ad- vances in neural information processing systems, 32

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.086531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:304673cab8642ebaa616175f7da6cbea0405ff52d04c58f1b24902af5888b9f6

Observation 00a87966-ab28-4326-8f84-b13c0990575b · outbound

This paper cites an unresolved cited work.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-17T19:40:10.222582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:5bd4b66f003b8d6f4f582d62d0135ccdd91cb4da327f34b27f2230e733ae6b38

Observation cad4028c-8d36-4655-8526-8593bc8c7ead · outbound

This paper cites GAReT: Cross-view Video Geolocalization with Adapters and Auto-Regressive Transformers.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale GAReT: Cross-view Video Geolocalization with Adapters and Auto-Regressive Transformers

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:02.554075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:6fc0683cdc077857cee0fb9d10be90e77aa973fb1cdac7d0fea5a57ade0643ff

Observation 4007da5c-8d6f-4e04-971f-3c55f3d308b6 · outbound

This paper cites Where in the world is this image? transformer-based geo-localization in the wild.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Where in the world is this image? transformer-based geo-localization in the wild

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.109949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:937e5e39b740e9c85ab22c1da3092a8f32d1ae45392e95c98fc0119392e3c656

Observation 15d0c818-5d3a-4283-8b34-a77ef687ffdd · outbound

This paper cites an unresolved cited work.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Unresolved cited work

Reference 38

Resolution
parse uncertain
raw_fallback, observed 2026-05-17T19:40:10.096338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:6d7467b591bae0656498c2dcbad9b4048a98bd28bd408b9206f9500c63a7dcdc

Observation a8a7f8a7-6a94-43a8-835c-a7daf8529522 · outbound

This paper cites Language models are unsu- pervised multitask learners.OpenAI blog, 1(8):9.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Language models are unsu- pervised multitask learners.OpenAI blog, 1(8):9

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.001694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:5af257241ee527af2d44d865fac6e9a1955e5897df09d5c2e25655c99ee350a4

Observation 5e9c9d57-4ce9-447c-a5b6-e22e3525e226 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Learning transferable visual models from natural language supervi- sion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.035218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:0d182d5a2071ba44c169c1f494b04fac5563ea16ec73913209e96aa51cc588cd

Observation 6346e42f-a180-4f8e-8bd5-1223262c1247 · outbound

This paper cites Cross-view image synthesis using geometry-guided conditional gans.Computer Vision and Image Understanding, 187:102788.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Cross-view image synthesis using geometry-guided conditional gans.Computer Vision and Image Understanding, 187:102788

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.021116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:5dc2443e6fcf2f43463a0148149f1da5750eadf163c057ba0a0a9814cdf1a40f

Observation ecbc984a-5752-4bba-875d-53922bbbd48b · outbound

This paper cites Bridging the domain gap for ground-to-aerial image matching.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Bridging the domain gap for ground-to-aerial image matching

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.129142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:0c106617d7cfc0c5489e61c40734675208b35147cf3aa34ab8597f14a0dc20b4

Observation c954a814-41de-47c3-aa98-1ba7a6adc2ce · outbound

This paper cites Video geo-localization employing geo-temporal feature learning and gps trajectory smoothing.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Video geo-localization employing geo-temporal feature learning and gps trajectory smoothing

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.152554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:02ceb2669384f5d30d6b267e66bf4b7d17f12b1d42a155bf8c63f81c840382a7

Observation 60b8ce06-2bd4-40ac-90f0-b0bbbf2f6c16 · outbound

This paper cites Generalized in- tersection over union: A metric and a loss for bounding box regression.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Generalized in- tersection over union: A metric and a loss for bounding box regression

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.038925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:865352cebee1501b9608359ab021291d296d0e09968224ea5db2e59fd5e946d9

Observation a83425b5-1d6f-415a-aa33-05099ddd9db5 · outbound

This paper cites The equal earth map projection.International Journal of Geographical Information Science, 33(3):454–465.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale The equal earth map projection.International Journal of Geographical Information Science, 33(3):454–465

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.106328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:8e4d856d3bc6d1b3cfc6385d10b9084c7d42413fb602c775c88697157f4b6c3d

Observation 7609aa96-2979-4323-96e7-1e6cda7e7f9f · outbound

This paper cites Cplanet: Enhancing image geolocalization by combi- natorial partitioning of maps.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Cplanet: Enhancing image geolocalization by combi- natorial partitioning of maps

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.064010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:08f09ec3bba239ceb2fcbdda514069b6e4d4fd2a224bda989de09e768846a994

Observation 9028f369-62f3-49d5-9da8-c3a6d1d582df · outbound

This paper cites Gt-loc: Unifying when and where in images through a joint embedding space.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Gt-loc: Unifying when and where in images through a joint embedding space

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.042196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:31f8acd7a2ab4b305dc7322b900aa46eb8743fbe7071d3e4f7243dc9d12ff577

Observation 823c50c6-4260-4fd3-9e01-908a4c89ef17 · outbound

This paper cites Where am i looking at? joint location and orientation es- timation by cross-view matching.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Where am i looking at? joint location and orientation es- timation by cross-view matching

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.214020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:d0778ff64185acf86a3356bbd4d47174bdb5932610fec05434a6cb869d86c289

Observation 9ca934de-886f-46a9-a3db-e6b8b9db1d1f · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:11:02.623353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:0bc2d1ad40f54ff1bc04bf3746c9169012cd5406b196c671e750ac74865a1d84

Observation 06d2f80a-f05e-4ef0-96b6-9c9916363398 · outbound

This paper cites Fourier features let networks learn high frequency functions in low dimen- 10 sional domains.Advances in neural information processing systems, 33:7537–7547.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Fourier features let networks learn high frequency functions in low dimen- 10 sional domains.Advances in neural information processing systems, 33:7537–7547

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.074386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:e23b8e8c4229484231e3530e673d6d15a637e62681c141f934e164fd2cf9f6b6

Observation 2d558be8-832b-4f80-992e-37ecacc0240a · outbound

This paper cites Coming down to earth: Satellite-to-street view synthesis for geo-localization.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Coming down to earth: Satellite-to-street view synthesis for geo-localization

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.218699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:4703388937ffc779bb82aa81b8ff1f32df9609e66672543d160c061ebfa2e248

Observation 8eee83f0-6d0f-486a-aab2-153a7765bd65 · outbound

This paper cites Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.Advances in neural information processing systems, 35:10078–10093.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training.Advances in neural information processing systems, 35:10078–10093

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.099804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:af2d3a189ce83a7eff2a14b4e2b2843349ddf738feca763c81c292b49e09ff1b

Observation b27accaf-d40c-48e5-8fa2-3f87ccdcb63e · outbound

This paper cites City scale geo-spatial trajectory estimation of a mov- ing camera.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale City scale geo-spatial trajectory estimation of a mov- ing camera

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.180087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:bb1c3a91bd6a849e66eb42f9268833a6727dba00e6d0e78fb4156cfc984c7eea

Observation a69a32f8-0f82-4aec-bbb6-37417e906c5b · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Attention is all you need.Advances in neural information processing systems, 30

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.191044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:4184418e572fc9735a548d1471c596fd7630b3934e78979cdc5f49e302e85258

Observation 4fe7ce10-1093-4f8f-bcb4-4bab0f93d6d0 · outbound

This paper cites Geoclip: Clip-inspired alignment be- tween locations and images for effective worldwide geo- localization.Advances in Neural Information Processing Systems, 36.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Geoclip: Clip-inspired alignment be- tween locations and images for effective worldwide geo- localization.Advances in Neural Information Processing Systems, 36

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.113453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:098e7897cd7b80fd2a63e399742c973449956e8ef9285d579ed86a2af82a93f4

Observation d9891760-cd4f-4c25-b0f1-cb89b3e8581b · outbound

This paper cites Revisiting im2gps in the deep learning era.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Revisiting im2gps in the deep learning era

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.017112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:ed7de8c9458b9794c203f8cc411609d7e39215090bc945cb227b95848fb48af4

Observation 48c0d0e8-887d-45b8-b7a4-da094bfa119f · outbound

This paper cites Gama: Cross- view video geo-localization.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Gama: Cross- view video geo-localization

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.149366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:5026083b43fcee56e3301e3682af17c8d1f0fa77b2790ff999fc28c97822f595

Observation b6d3b66e-279a-44ea-ab44-187aabe4edd5 · outbound

This paper cites fairseq S2T: Fast Speech-to-Text Modeling with fairseq.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale fairseq S2T: Fast Speech-to-Text Modeling with fairseq

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:02.475816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:ca2f7b00f771b8e13208249aa76699384cc00c2bfc68832ed7fce696f19022ec

Observation 2c1ae000-a86e-4095-a839-ba23421e9284 · outbound

This paper cites Mapillary street-level sequences: A dataset for lifelong place recognition.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Mapillary street-level sequences: A dataset for lifelong place recognition

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.145808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:6f9cc03846f75f0434113247414a852ce64c75130a28899d6e46890beece9006

Observation 1f7fb1da-ddf5-405a-9384-836c28c90bb7 · outbound

This paper cites Planet- photo geolocation with convolutional neural networks.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Planet- photo geolocation with convolutional neural networks

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:09.986691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:9cdb957c0c658987c0b4d01c07819f48910db7173eccec24461eb18d0e7db56f

Observation c0ef8d27-5505-4395-bbfc-7781fa5e8317 · outbound

This paper cites Wide-area image geolocalization with aerial reference im- agery.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Wide-area image geolocalization with aerial reference im- agery

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.226812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:ea068c23f6f5014a3b7e17a92a2d66f90e1f8426249fae5bc045b2dafa8b2d13

Observation cafbd990-faf6-46d0-be3d-7b42a12adb31 · outbound

This paper cites Cross-view geo-localization with layer-to-layer transformer.Advances in Neural Information Processing Systems, 34:29009–29020.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Cross-view geo-localization with layer-to-layer transformer.Advances in Neural Information Processing Systems, 34:29009–29020

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.117156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:979fc21b8780ddef603f21968add0c28318fb37c625f038b7853a4a56139c097

Observation b93c00ac-06f6-40fb-b1b1-fc65248d2931 · outbound

This paper cites Bdd100k: A diverse driving dataset for heterogeneous multitask learning.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Bdd100k: A diverse driving dataset for heterogeneous multitask learning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.123217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:99feb603b74c11340316b11674779f38282c28957dc025d91223b4748f5cde43

Observation a21ee39c-aa93-4b00-ab25-1d476df4c22f · outbound

This paper cites Sigmoid loss for language image pre-training.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Sigmoid loss for language image pre-training

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:09.991174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:13be641184d2335fdc9dfbdfddbb57dc87ecc15da16bb9daf42e8449cf1cfd3e

Observation 7a8910d8-cbdc-4c87-a9d2-d677adfc32b5 · outbound

This paper cites Places: A 10 million image database for scene recognition.IEEE transactions on pattern analysis and machine intelligence, 40(6):1452–1464.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Places: A 10 million image database for scene recognition.IEEE transactions on pattern analysis and machine intelligence, 40(6):1452–1464

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.025548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:a128c0d34998322d49eda065941fb6e1ee391478fe3a3c506ea3ef60c96a3d56

Observation 0bb0feb4-d82f-456f-92de-d1488567abda · outbound

This paper cites Vigor: Cross- view image geo-localization beyond one-to-one retrieval.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Vigor: Cross- view image geo-localization beyond one-to-one retrieval

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.132450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:89d6c0faea8b2ce1f12cdc146cebc9a7416a43987977943ab0f0444d2ac7c8ec

Observation cbc85809-bd1f-42a7-a566-30467dc98112 · outbound

This paper cites Transgeo: Trans- former is all you need for cross-view image geo-localization.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Transgeo: Trans- former is all you need for cross-view image geo-localization

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.012305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:33478f23ae64d1b7b97d082e966df08cc92629456e4771781a2b2db450dc84b6

Observation 92610c0a-e5e4-4013-a3b9-7c1106e3914d · outbound

This paper cites (If there are a few outliers they can be skipped).

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale (If there are a few outliers they can be skipped)

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.155922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:9a05749855bc2ab2f29859adb7a05c17f2c1f7afe62e28ba8a66125d1154e54d

Observation d199c9e3-30a7-4208-804d-29887df72c1e · outbound

This paper cites Also determine the resolution of the gallery (the finer the resolution the larger the gallery.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Also determine the resolution of the gallery (the finer the resolution the larger the gallery

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.061056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:5250d80601e08b9d1a9ea2b1b6c89da4207a7af91dc4efd89c4c80329b741e35

Observation 9cbe8649-66c7-43f7-a667-ea0f06c67315 · outbound

This paper cites an unresolved cited work.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-05-17T19:40:10.163421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:ca2e439a058c4dff1fb040b399951bfc579b3b97bc64284a828472db7344b13b

Observation c435deef-c37f-4bf1-86d3-45a90f1bd3f6 · outbound

This paper cites an unresolved cited work.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-05-17T19:40:10.051430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:fc39951e879537315215d7e2503b4debb1e0f5aebea2a2b51a36835681ba7126

Observation 195641ed-dd49-40fc-b644-3b5ea8d5dc7e · outbound

This paper cites an unresolved cited work.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-05-17T19:40:10.048329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:26790e5ab05c04cbb078181f892826978f051bad2baf81f2d2673ae3ec2c2c6d

Observation 136c5760-b6e2-4eb2-bf6e-119bbc178dde · outbound

This paper cites an unresolved cited work.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-05-17T19:40:10.057871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:386a9b4fead95391a94e52640640d526a1e875963c081c9630c589c28e7ad413

Observation 0cad1e0b-cc1a-40c9-8486-34ea4535f529 · outbound

This paper cites This will serve as the ground truth label of the entire video sequence.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale This will serve as the ground truth label of the entire video sequence

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.008356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:55e2e8154d973acdf6262faebc827fcd37033825f7e8eb52c5be06dc541faf6d

Observation 15fbe9a9-3b93-4cd1-90d2-6b31552ebd1d · outbound

This paper cites Obtain the prediction that is closest to this centroid (this reduces error due to out- liers).

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Obtain the prediction that is closest to this centroid (this reduces error due to out- liers)

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.159738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:6088011bfde8dd47aa24dc351c93004995d134975677913a8a8c94c6c3c19cbd

Observation 8a51bb11-ed4c-4ee0-87ab-147c1f016444 · outbound

This paper cites Compute distance accuracy at all thresholds using this distance measure.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Compute distance accuracy at all thresholds using this distance measure

Reference 76

Resolution
malformed identifier
raw_fallback, observed 2026-05-17T19:40:10.070759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:e32a088117634419924141d6a6628a1637924a10bea84bdf4f9a7a2b6a5a1cad

Observation ba5aa688-d127-4f94-847f-84ac4469a9fd · outbound

This paper cites an unresolved cited work.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-05-17T19:40:10.077000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:5948f2b672270c2e0b47caa0d1512d83554df9e3833045e25edcccd893d57329

Observation 223f54a2-361b-4743-8715-79ccff5c4253 · outbound

This paper cites (b) As CityGuessr68k is very large, we sample the data in such a way that the number of sequences is roughly equivalent to MSLS.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale (b) As CityGuessr68k is very large, we sample the data in such a way that the number of sequences is roughly equivalent to MSLS

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:09.998205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:9cef2dacbb7186a151d473ec0774be72cebf53107448924adccee466db9982c7

Observation ba8e19a2-9419-4371-806c-7e885756aa3b · outbound

This paper cites We train a model on this unified data for 200 epochs at an learning rate decay rate of 0.97.

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale We train a model on this unified data for 200 epochs at an learning rate decay rate of 0.97

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T19:40:10.004948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:37:16.031502Z digest=sha256:082b9615f68c9b20d2de20970535ca1b9e00093998ef013e23442ca7ccca5d78

Pith citing papers

No inbound Pith citation observations are available.