Pith. sign in

Paper Citation Record · LEDGER

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization

As of 20 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2507.04509.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04509 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:51:53.074321Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 638fde49-e415-43b4-bea2-c236ad418035 · outbound

This paper cites Random sample consensus: A paradigm for model fitting with applications to image analysis and automated cartography,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Random sample consensus: A paradigm for model fitting with applications to image analysis and automated cartography,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:58.329827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:50.260692Z digest=sha256:3062a39d2d66e7ce1f7776b278e0581b2c41bfc220a4a623f2e18f5c812176f0

Observation 362dd15c-f339-4545-95e0-e0491b5f214d · outbound

This paper cites Posenet: A convolutional network for real-time 6-dof camera relocalization,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Posenet: A convolutional network for real-time 6-dof camera relocalization,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:58.175786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:50.355729Z digest=sha256:9b46f34fe4bf021be26eb0bc8e686a8e29eb02681ddb5a45509b3a9c1ec92c21

Observation 6e56b12d-7ddd-47d2-8c99-9342ec23b794 · outbound

This paper cites Image-based localization using lstms for structured feature correlation,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Image-based localization using lstms for structured feature correlation,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:58.050472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:50.425621Z digest=sha256:db9edd93be8aa9405fd98bdc98d08cd2af5a95af47c79e86960bae1678a60530

Observation d23f89fb-7306-462d-b63d-0850a8f6cd66 · outbound

This paper cites Image-based localization using hourglass networks,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Image-based localization using hourglass networks,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:57.962733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:50.505453Z digest=sha256:6144f122378af8b43eee99a67b88a2c824b97aded4e0768db80feae36434b5f5

Observation 96a5534c-f03e-4db4-9c85-9049bf354677 · outbound

This paper cites Modelling uncertainty in deep learning for camera relocalization,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Modelling uncertainty in deep learning for camera relocalization,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:57.814913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:50.541883Z digest=sha256:f2e9e3490fc2cc32719a020b5a10fbe8470b2d60e926d3d6149860bd6570b497

Observation 7d8c9377-86e9-4e77-b200-acd37f6af818 · outbound

This paper cites Geometric loss functions for camera pose regression with deep learning,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Geometric loss functions for camera pose regression with deep learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:57.672985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:50.649083Z digest=sha256:6e3f293228a30e00eb5fd0001990291c117caef73f701f88f16e39d53950dc8b

Observation f5555b2d-9c3e-4b87-825d-3a101907eee6 · outbound

This paper cites Vidloc: A deep spatio-temporal model for 6-dof video-clip relocalization,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Vidloc: A deep spatio-temporal model for 6-dof video-clip relocalization,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:57.529663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:50.708062Z digest=sha256:fbbe05b660b5a07071b6b6c52abf4bc2dfc7720e354407c096db0a66198ff211

Observation c35f19f2-66a0-4abc-b7f7-82da0b72176c · outbound

This paper cites Atloc: Attention guided camera localization,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Atloc: Attention guided camera localization,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:57.416565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:50.797015Z digest=sha256:c8e6ae863f991b6274f0fd1db69f6be59306ac637fd903ed80d3435d6135e6bf

Observation 91c298e2-af3a-43ba-8ed9-cfd3119541c6 · outbound

This paper cites Effloc: Lightweight vision transformer for effi- cient 6-dof camera relocalization,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Effloc: Lightweight vision transformer for effi- cient 6-dof camera relocalization,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:57.300384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:50.878081Z digest=sha256:16d18de82badef762b440a9392d181afcd6d1d966f23cd5abea791bfade7b6fd

Observation 11c3bad6-0b8d-4f07-b6de-e5207a512cf4 · outbound

This paper cites NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:51:50.963143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:51:50.963143Z digest=sha256:ef220994c8733f6d3b7ca72fe7adfeb72e35616016707d1798bf8f879bafdb47

Observation 8a0113dc-d0f3-4fa9-97b9-53026008ccf0 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:51:51.065164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:51:51.065164Z digest=sha256:686918bb618d542080221bc9dcc078c12e336333bbb173afcb8d085c3e41ee0f

Observation ef2c3e76-431d-4d3d-a9c4-372d8b0ea5f7 · outbound

This paper cites Extending absolute pose regression to multiple scenes,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Extending absolute pose regression to multiple scenes,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:57.175742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:51.145193Z digest=sha256:7dfcb4f0b70c4803c6fec6b6137ba78462d207f2d22f70f770684ec45181d15c

Observation 3c40f7ac-7a35-4f5d-ba29-539d8f14ddcd · outbound

This paper cites Coarse-to-fine multi-scene pose regression with trans- formers,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Coarse-to-fine multi-scene pose regression with trans- formers,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:57.053636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:51.218239Z digest=sha256:92e757cbb26481ea4c47ebe0deb6fe9842dfaf552e86ef53c0438469642824bd

Observation 6df5821e-0bac-4bdf-86b9-8ac1b311214a · outbound

This paper cites City-scale landmark identification on mobile devices,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization City-scale landmark identification on mobile devices,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:56.892027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:51.307081Z digest=sha256:e6e7b521f51b9c116da9aeb876a623666f83f43226d551f9fafdbfb3f573df03

Observation c13af43b-2b0c-4343-b543-574fb85d132c · outbound

This paper cites Imagdressing-v1: Cus- tomizable virtual dressing,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Imagdressing-v1: Cus- tomizable virtual dressing,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:56.661234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:51.394012Z digest=sha256:37ed9d4d2698a854ad5504076a419cd57023ddd7c24b4110dc7f1fba0d04b12c

Observation 3243608f-e3ab-4310-9934-8b36b0961e8b · outbound

This paper cites Imagpose: A unified conditional framework for pose-guided person generation,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Imagpose: A unified conditional framework for pose-guided person generation,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:51:51.465976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:51:51.465976Z digest=sha256:06df411e494e5828737aa6f1177afd8c7e0d83b12fac6f17ac27e854782e5d70

Observation 28313878-7355-4dd8-88e7-2d146d3d4dd9 · outbound

This paper cites Dsac-differentiable ransac for camera localization,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Dsac-differentiable ransac for camera localization,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:56.366071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:51.567426Z digest=sha256:94fd2d8164512c31bddc677872eac22abafa35eedc0080d8c4348cb06b3c1df5

Observation eaf78365-3f39-4339-94d2-fb681fe3373e · outbound

This paper cites Hybrid scene compression for visual localization,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Hybrid scene compression for visual localization,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:56.099167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:51.604086Z digest=sha256:6a77824559980afc78f44fe4e738cc7d17049127979d2084329ea7923622cfb0

Observation c0b92286-dabf-42d7-a745-590ecbcb0bc2 · outbound

This paper cites Map-free visual relocalization: Metric pose relative to a single image,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Map-free visual relocalization: Metric pose relative to a single image,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:55.912782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:51.641587Z digest=sha256:8c62a9e90d2b1345abca059efa2b41ce84ec44c3c41cd11d2e2bb237b7e767df

Observation cec9be87-48db-4302-8628-d1ced13e8a31 · outbound

This paper cites Learning multi-scene absolute pose regression with transformers,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Learning multi-scene absolute pose regression with transformers,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:55.740171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:51.742970Z digest=sha256:7766446c51e2e14ec8e656d20c08821a035bc051ca4348fd5c9e5a98ea6bfe35

Observation cc0e3ccf-ea5f-405f-b71a-94df2dabad27 · outbound

This paper cites Geometry-aware learning of maps for camera localization,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Geometry-aware learning of maps for camera localization,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:55.618580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:51.798687Z digest=sha256:c444f3ae28320ca449a158b70d5224a944051fadb07dadffd6da19c4e5bb6727

Observation c44a2cfe-b1cc-4474-941d-dcaf2fd101b9 · outbound

This paper cites Fusionloc: Camera-2d lidar fusion using multi-head self-attention for end-to-end serving robot relocalization,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Fusionloc: Camera-2d lidar fusion using multi-head self-attention for end-to-end serving robot relocalization,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:55.452563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:51.894716Z digest=sha256:2235bea3c762d9542fae43f2938300efe09daae584452172a31e038c2cda23de

Observation 75f786dc-6c23-417d-8104-ed146e475ff7 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Learning transferable visual models from natural language supervision,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:55.249451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:51.949874Z digest=sha256:6baa27147796e39d7670196f32e467e04ca46114551faeec77bf8dc86a1ad392

Observation 7b745bcb-5c88-475d-b6fc-c15536b35469 · outbound

This paper cites CLIP-Adapter: Better Vision-Language Models with Feature Adapters.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization CLIP-Adapter: Better Vision-Language Models with Feature Adapters

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:51:52.021631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:51:52.021631Z digest=sha256:970a1de2667262c0a61fad295c4fcb0aed3f18e982c9267d54690f0bcd13785d

Observation e814071b-bf23-4026-8066-692a376fc4c7 · outbound

This paper cites Envedit: Environment editing for vision-and-language naviga- tion,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Envedit: Environment editing for vision-and-language naviga- tion,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:55.073543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:52.093706Z digest=sha256:369a37fa705c9ef40aefc92b42c5881dd191f37ff1bfbde985b9534a991ec5e6

Observation eb68e5e1-1336-4a25-be70-1f42839c86d7 · outbound

This paper cites Learning to generate scene graph from natural language supervision,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Learning to generate scene graph from natural language supervision,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:54.966464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:52.204735Z digest=sha256:143bb6ae5a209a6d70807266379082ac0a6e19c5ebfc04e48f8ed60988432e3b

Observation c666e258-ce83-4cdc-87b0-d6f1eb17bad3 · outbound

This paper cites Denseclip: Language-guided dense prediction with context-aware prompting,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Denseclip: Language-guided dense prediction with context-aware prompting,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:54.855050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:52.282022Z digest=sha256:2d1c345528abacb9f55922a246573690120ba712d9616c4f9767ef0a0c70f569

Observation 9afee136-979f-4461-ad99-dda00c0805d6 · outbound

This paper cites Fm-loc: Using foundation models for improved vision-based localization,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Fm-loc: Using foundation models for improved vision-based localization,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:54.708146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:52.397707Z digest=sha256:9ca2b35fd6d012362dac5e9fa97b9e9a9347dce3ec13521fcbd6323ce424c5cf

Observation b37d3e6a-e5fd-4118-b89b-61be37dd36e0 · outbound

This paper cites Clip-loc: Multi-modal landmark association for global localization in object-based maps,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Clip-loc: Multi-modal landmark association for global localization in object-based maps,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:54.488600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:52.488586Z digest=sha256:8e0494745d217ea40c8a93d0a28defaf0b5b14851b350b7409ad8e5906189e7c

Observation f1622f84-0953-49c4-8d5a-4610bba884a6 · outbound

This paper cites Geollm: Extracting geospatial knowledge from large language models,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Geollm: Extracting geospatial knowledge from large language models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:54.312406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:52.561556Z digest=sha256:c78b7fb27a533d2e17ca17a30584a0930d83ff84e4302c6dcba085f935e37db9

Observation 40438513-e33f-406f-9f44-371719b0664f · outbound

This paper cites Georeasoner: Geo-localization with reasoning in street views using a large vision-language model,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Georeasoner: Geo-localization with reasoning in street views using a large vision-language model,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:54.112020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:52.660280Z digest=sha256:f0ca31ac8e18d68a5c029549e0af5ccdcc41ce39e70283eca08659ec2b188d4c

Observation 914a639e-3631-4c88-8f43-7a73f31c2c34 · outbound

This paper cites Advancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion Models.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Advancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:51:52.756841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:51:52.756841Z digest=sha256:2b6df155c394ef0a15ef2d681f08f42023fe24ac8d33f6e86f96125c58f27762

Observation 64f9f7c1-362b-477c-a6a6-2db52a0c705f · outbound

This paper cites Boosting consistency in story visualization with rich-contextual conditional diffusion models,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Boosting consistency in story visualization with rich-contextual conditional diffusion models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:53.933474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:52.847978Z digest=sha256:5ec9539142990416396c994d1991a11057d2b6987391221583ea6e5041ef0190

Observation 2e3092c5-4a38-4a7d-a651-017588757960 · outbound

This paper cites 7-scenes dataset,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization 7-scenes dataset,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:53.706874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:52.928410Z digest=sha256:47be787c72838c4a2ede73abf6c19ec0189bbbb0ceb8e8ee96b5b3bc6bde8d91

Observation cbb96d11-fce4-4cb5-9e62-0dc5cccb8b9c · outbound

This paper cites Do we really need scene-specific pose encoders?.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Do we really need scene-specific pose encoders?

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:53.521611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:53.005097Z digest=sha256:c9c9872b83803fd2097598e9d7a65d4b94aeb09cbfaa0421c3cface1643ad7f4

Observation d4c93214-bca1-4038-99db-956f9694ab88 · outbound

This paper cites Image-based localization using hourglass networks,.

MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization Image-based localization using hourglass networks,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:51:53.328502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T19:51:53.074321Z digest=sha256:569a43087b9bdba51959c49bb1920f348bfafa0395e9011988214065361d2771

Pith citing papers

No inbound Pith citation observations are available.