Pith. sign in

Paper Citation Record · LEDGER

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization

As of 7 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2507.04509.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04509 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:51:53.074321Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 638fde49-e415-43b4-bea2-c236ad418035 · outbound

This paper cites Random sample consensus: A paradigm for model fitting with applications to image analysis and automated cartography,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Random sample consensus: A paradigm for model fitting with applications to image analysis and automated cartography,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:58.329827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:50.260692Z digest=sha256:622b6854b7898c75f20587cbcb393cc6696fe7eb9f131051272e7f2a86bed611

Observation 362dd15c-f339-4545-95e0-e0491b5f214d · outbound

This paper cites Posenet: A convolutional network for real-time 6-dof camera relocalization,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Posenet: A convolutional network for real-time 6-dof camera relocalization,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:58.175786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:50.355729Z digest=sha256:2f55228d48afd9249755cfe968605b6e75a849548c599660f3c1e5a5b4c753bd

Observation 6e56b12d-7ddd-47d2-8c99-9342ec23b794 · outbound

This paper cites Image-based localization using lstms for structured feature correlation,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Image-based localization using lstms for structured feature correlation,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:58.050472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:50.425621Z digest=sha256:89bfd4800ebc9a931870b733d15c8dc49b577b62e4d790e7a816126532012c05

Observation d23f89fb-7306-462d-b63d-0850a8f6cd66 · outbound

This paper cites Image-based localization using hourglass networks,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Image-based localization using hourglass networks,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:57.962733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:50.505453Z digest=sha256:09a8263c0c9cc1600246c3c45b5369263eb1269032961678ca8b0704a6334950

Observation 96a5534c-f03e-4db4-9c85-9049bf354677 · outbound

This paper cites Modelling uncertainty in deep learning for camera relocalization,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Modelling uncertainty in deep learning for camera relocalization,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:57.814913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:50.541883Z digest=sha256:eb2a98779d039ae40a4a5d8a1a852e861d2df8f3a4780f28d23e0164403f9e47

Observation 7d8c9377-86e9-4e77-b200-acd37f6af818 · outbound

This paper cites Geometric loss functions for camera pose regression with deep learning,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Geometric loss functions for camera pose regression with deep learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:57.672985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:50.649083Z digest=sha256:9b9bc52fbd2b8923be1db861686568141668e7c59f72d49d4feb4dc2919f9556

Observation f5555b2d-9c3e-4b87-825d-3a101907eee6 · outbound

This paper cites Vidloc: A deep spatio-temporal model for 6-dof video-clip relocalization,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Vidloc: A deep spatio-temporal model for 6-dof video-clip relocalization,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:57.529663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:50.708062Z digest=sha256:4ab2445cb38cc16ddd7ec1f63052a5a5ba7716404548c7fd512186acf651d617

Observation c35f19f2-66a0-4abc-b7f7-82da0b72176c · outbound

This paper cites Atloc: Attention guided camera localization,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Atloc: Attention guided camera localization,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:57.416565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:50.797015Z digest=sha256:3711bfafa15ab6ee618e07aa8ca31e2410e20196819203d2feece325b79cd402

Observation 91c298e2-af3a-43ba-8ed9-cfd3119541c6 · outbound

This paper cites Effloc: Lightweight vision transformer for effi- cient 6-dof camera relocalization,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Effloc: Lightweight vision transformer for effi- cient 6-dof camera relocalization,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:57.300384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:50.878081Z digest=sha256:5774424ea0e3d3093b40d90f351dd90c41017130cdf3a97f9351049548107129

Observation 11c3bad6-0b8d-4f07-b6de-e5207a512cf4 · outbound

This paper cites NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:51:50.963143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:51:50.963143Z digest=sha256:f8f0898c56b97868054ef24690abdcc207d72e8b81a095cd1f5a55e1cf3d63db

Observation 8a0113dc-d0f3-4fa9-97b9-53026008ccf0 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:51:51.065164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:51:51.065164Z digest=sha256:66fcf5f1e28e079dc7663bbc05cf8f664038f2d84a496e2709bb91497a69ec57

Observation ef2c3e76-431d-4d3d-a9c4-372d8b0ea5f7 · outbound

This paper cites Extending absolute pose regression to multiple scenes,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Extending absolute pose regression to multiple scenes,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:57.175742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:51.145193Z digest=sha256:3e0485c438af2e893719c7fe5b09dea9c5d6ec7b604b1b51efd087593f76b0a8

Observation 3c40f7ac-7a35-4f5d-ba29-539d8f14ddcd · outbound

This paper cites Coarse-to-fine multi-scene pose regression with trans- formers,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Coarse-to-fine multi-scene pose regression with trans- formers,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:57.053636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:51.218239Z digest=sha256:5fe923690e056a3b667f66de0dc7cd8d6c5f2d98e15aab12570221e446600a5f

Observation 6df5821e-0bac-4bdf-86b9-8ac1b311214a · outbound

This paper cites City-scale landmark identification on mobile devices,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization City-scale landmark identification on mobile devices,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:56.892027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:51.307081Z digest=sha256:dafc7f9974a6a17709c09a609b8d07d750e86d71b03cfa1bbc692864876e5319

Observation c13af43b-2b0c-4343-b543-574fb85d132c · outbound

This paper cites Imagdressing-v1: Cus- tomizable virtual dressing,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Imagdressing-v1: Cus- tomizable virtual dressing,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:56.661234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:51.394012Z digest=sha256:5e55aee92ad1ba604492a9abbef8873c9666dc87083ab58eca28d1f51868a82f

Observation 3243608f-e3ab-4310-9934-8b36b0961e8b · outbound

This paper cites Imagpose: A unified conditional framework for pose-guided person generation,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Imagpose: A unified conditional framework for pose-guided person generation,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:51:51.465976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:51:51.465976Z digest=sha256:c5e6d447c5e2f7d1afdfbd2c54e1d441696bb6fd249c9f76d6445f0c2cd471f0

Observation 28313878-7355-4dd8-88e7-2d146d3d4dd9 · outbound

This paper cites Dsac-differentiable ransac for camera localization,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Dsac-differentiable ransac for camera localization,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:56.366071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:51.567426Z digest=sha256:1c1ea89102c8bb1605af8cfea4c6472eedb4bdd368e939e0c83ed207ff0b257c

Observation eaf78365-3f39-4339-94d2-fb681fe3373e · outbound

This paper cites Hybrid scene compression for visual localization,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Hybrid scene compression for visual localization,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:56.099167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:51.604086Z digest=sha256:eb519f075548494df6860cbc45a4b331f943854a5a71a49c99dcb4ccdb2ba255

Observation c0b92286-dabf-42d7-a745-590ecbcb0bc2 · outbound

This paper cites Map-free visual relocalization: Metric pose relative to a single image,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Map-free visual relocalization: Metric pose relative to a single image,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:55.912782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:51.641587Z digest=sha256:f1fa7fc1c621368efd10b0bdd1b0f39d6c50fcb405729be0964b201a110ef91f

Observation cec9be87-48db-4302-8628-d1ced13e8a31 · outbound

This paper cites Learning multi-scene absolute pose regression with transformers,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Learning multi-scene absolute pose regression with transformers,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:55.740171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:51.742970Z digest=sha256:844eb522a95239dd47587a0e261b7d86095cd0408f933777c3ce2328b75a4441

Observation cc0e3ccf-ea5f-405f-b71a-94df2dabad27 · outbound

This paper cites Geometry-aware learning of maps for camera localization,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Geometry-aware learning of maps for camera localization,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:55.618580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:51.798687Z digest=sha256:3c05eb26ead64b917c8b0ceea7cac2f43fcc9aff02a722714bd38271bd4e64e3

Observation c44a2cfe-b1cc-4474-941d-dcaf2fd101b9 · outbound

This paper cites Fusionloc: Camera-2d lidar fusion using multi-head self-attention for end-to-end serving robot relocalization,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Fusionloc: Camera-2d lidar fusion using multi-head self-attention for end-to-end serving robot relocalization,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:55.452563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:51.894716Z digest=sha256:2e2d7073a26e6f9cbb20c2f3c8ba81eb88e10266e31ecc48e9fb65f7ca1afbfa

Observation 75f786dc-6c23-417d-8104-ed146e475ff7 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Learning transferable visual models from natural language supervision,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:55.249451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:51.949874Z digest=sha256:958b9bc6ed6248cfb8c9a60dbfb5a2fdf3ce342150886fae2afd7bc12a78df9c

Observation 7b745bcb-5c88-475d-b6fc-c15536b35469 · outbound

This paper cites CLIP-Adapter: Better Vision-Language Models with Feature Adapters.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:51:52.021631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:51:52.021631Z digest=sha256:80786a41578017e75e0731aaec3c71dea1b4cabab37507d52c5b1c4acc1255a3

Observation e814071b-bf23-4026-8066-692a376fc4c7 · outbound

This paper cites Envedit: Environment editing for vision-and-language naviga- tion,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Envedit: Environment editing for vision-and-language naviga- tion,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:55.073543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:52.093706Z digest=sha256:7662b3d49949db5af96d3c10c3c977f9f33716140f7ed0d0b8c538583c6db1e2

Observation eb68e5e1-1336-4a25-be70-1f42839c86d7 · outbound

This paper cites Learning to generate scene graph from natural language supervision,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Learning to generate scene graph from natural language supervision,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:54.966464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:52.204735Z digest=sha256:a267b24703aacd2c7a68fdf348ce9ec162027981b7e704c8b861e594b3ada22e

Observation c666e258-ce83-4cdc-87b0-d6f1eb17bad3 · outbound

This paper cites Denseclip: Language-guided dense prediction with context-aware prompting,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Denseclip: Language-guided dense prediction with context-aware prompting,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:54.855050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:52.282022Z digest=sha256:ea665d0ce74e2142da6dbf070295a40d282a4e204e5b2e07339656c0d8ef7fcd

Observation 9afee136-979f-4461-ad99-dda00c0805d6 · outbound

This paper cites Fm-loc: Using foundation models for improved vision-based localization,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Fm-loc: Using foundation models for improved vision-based localization,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:54.708146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:52.397707Z digest=sha256:ce3650be94393f439e72df9340e7aae84c2f075ff8d6c11fcf8204ee4515d41c

Observation b37d3e6a-e5fd-4118-b89b-61be37dd36e0 · outbound

This paper cites Clip-loc: Multi-modal landmark association for global localization in object-based maps,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Clip-loc: Multi-modal landmark association for global localization in object-based maps,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:54.488600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:52.488586Z digest=sha256:a9e11e7170616d2ac71b5de74b27e0e97329f6e6bb3e78db1bef3a716f3aeefd

Observation f1622f84-0953-49c4-8d5a-4610bba884a6 · outbound

This paper cites Geollm: Extracting geospatial knowledge from large language models,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Geollm: Extracting geospatial knowledge from large language models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:54.312406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:52.561556Z digest=sha256:aff586492019587734e75c0b003cc94b6c4ac1ba44833ed468a2d7abd728dd3b

Observation 40438513-e33f-406f-9f44-371719b0664f · outbound

This paper cites Georeasoner: Geo-localization with reasoning in street views using a large vision-language model,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Georeasoner: Geo-localization with reasoning in street views using a large vision-language model,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:54.112020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:52.660280Z digest=sha256:3b7bb51d6930e5c26c91ab330da7133df10cba4fb4fabeaa030ea22c849238d2

Observation 914a639e-3631-4c88-8f43-7a73f31c2c34 · outbound

This paper cites Advancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion Models.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Advancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:51:52.756841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:51:52.756841Z digest=sha256:51b1dc1ac34e86d289cee42bd0111b016ad45d8592d75ef312a12755bcabcc4d

Observation 64f9f7c1-362b-477c-a6a6-2db52a0c705f · outbound

This paper cites Boosting consistency in story visualization with rich-contextual conditional diffusion models,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Boosting consistency in story visualization with rich-contextual conditional diffusion models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:53.933474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:52.847978Z digest=sha256:5432e9064c36e7ba51da6c244d27b9fb33c2cbd5ad3a9e949da421b5afcb4866

Observation 2e3092c5-4a38-4a7d-a651-017588757960 · outbound

This paper cites 7-scenes dataset,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization 7-scenes dataset,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:53.706874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:52.928410Z digest=sha256:12bced3e2a50d50a9207a6caba5dfc070a604781ce0f8e62899df44513e16461

Observation cbb96d11-fce4-4cb5-9e62-0dc5cccb8b9c · outbound

This paper cites Do we really need scene-specific pose encoders?.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Do we really need scene-specific pose encoders?

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:53.521611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:53.005097Z digest=sha256:ed9547a44f6ad95fdcc4d9d61d6fa77dab866312635ea8d7a2b8175a143b9cb2

Observation d4c93214-bca1-4038-99db-956f9694ab88 · outbound

This paper cites Image-based localization using hourglass networks,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Image-based localization using hourglass networks,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:53.328502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:51:53.074321Z digest=sha256:e75cfe0d3b399c5d969396d1beecbbb0a90ecb598e1b05b9027600c6fefad2ef

Pith citing papers

No inbound Pith citation observations are available.