Pith. sign in

Paper Citation Record · LEDGER

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields

As of 7 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 0 inbound Pith citation observations for arXiv:2506.23352.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23352 v1

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:49:48.230001Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

96 of 96 outbound references displayed

  • verified exact2
  • verified fuzzy64
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9a2a456f-48b1-461b-9bab-b5febf5309c0 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Scanqa: 3d question answering for spatial scene understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.054070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.054070Z digest=sha256:431fc4f78c0c1bec118094b38f311553ea22d3ca21712a39094c755dae21037f

Observation 7bba69c6-6d78-4fd5-89d6-9cdd55d2ada1 · outbound

This paper cites Qwen2.5-VL Technical Report.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.094181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.094181Z digest=sha256:371f13eff48adc03578c0f1bcf6e5d24b3ca6fbde5c428c6887b5b9bb4a1f39a

Observation 3922716b-968a-46df-9b06-41652b632605 · outbound

This paper cites Henriques, Andrew Zisserman, and Andrea Vedaldi.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Henriques, Andrew Zisserman, and Andrea Vedaldi

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.194838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.194838Z digest=sha256:af66eb3459674406db35a04732f3b865710828e87da503daf75d8c5920122736

Observation 05a9f140-72af-403e-8111-ef8f7d6c539b · outbound

This paper cites Blumer, Qingx- uan Chen, and Francis Engelmann.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Blumer, Qingx- uan Chen, and Francis Engelmann

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.256991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.256991Z digest=sha256:d8f2ab5218f0af4820fe98b94c0c58c3fe5740d32a4fad39fcb1cf8e6b189238

Observation b093582a-470c-4d28-b589-ce745a4c8e3c · outbound

This paper cites A persistent spatial semantic representation for high-level natural language instruction execution.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields A persistent spatial semantic representation for high-level natural language instruction execution

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.392137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.392137Z digest=sha256:e8a2c39d46069e9951c9be2f0fe4ec6ae9c72d2c1c9d2d69a01996b1a789b20a

Observation 71ca8a59-5ff8-47a7-9292-8244e40f89d9 · outbound

This paper cites Prompt-rsvqa: Prompt- ing visual context to a language model for remote sensing visual question answering.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Prompt-rsvqa: Prompt- ing visual context to a language model for remote sensing visual question answering

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.499765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.499765Z digest=sha256:89a976797569508eb9f650c70bd3bba02acacd3490444b5c092af8bf3eb80d7b

Observation dee5502b-394d-4b8a-9d94-53bcfd086c67 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natu- ral language.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Scanrefer: 3d object localization in rgb-d scans using natu- ral language

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.634508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.634508Z digest=sha256:8e7a281a8fcb0f3d014554a38cec2e54d677fa9a0993a6d9e467224046b12ce2

Observation 2d18a5ce-27c7-4499-9d5a-6992720f8660 · outbound

This paper cites Panoptic vision-language feature fields.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Panoptic vision-language feature fields

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.745471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.745471Z digest=sha256:c3b9715f49bf28819c45e252cec16401c5b957f7eb0ad7a741acecda753cadd8

Observation e45035e5-1c0c-469c-a16e-8289272ccb08 · outbound

This paper cites Stylecity: Large-scale 3d urban scenes stylization.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Stylecity: Large-scale 3d urban scenes stylization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.839057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.839057Z digest=sha256:2846162d79b15a0dca88796d474af4b62c2570cb967f054c86419061a52714a5

Observation d71297e6-2235-430c-96e6-bf273376ff1e · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.929873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.929873Z digest=sha256:a4bd542357da8458a4814152af38082c8a343255d857ffa7fd5e6ba72467c85b

Observation 00fe7b97-3e74-4516-ab3c-9da16c510e9c · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.012679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.012679Z digest=sha256:52cc3a3b111d716ff1bb186d09043bc2a41635ec1f8dd29aae70803d800948a7

Observation da1aa099-20a3-46e0-a396-41eb294e5e9e · outbound

This paper cites LayoutGPT: Compositional Visual Planning and Generation with Large Language Models.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields LayoutGPT: Compositional Visual Planning and Generation with Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.132312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.132312Z digest=sha256:8f6e554b00dffae344e6bab6b7aaf8dca1d8c526ee5d4ed03a0ed4fa7da48cfc

Observation 1999571b-cd9a-407c-bc5f-ccfc1b13af7e · outbound

This paper cites Dynamic 3d gaussian fields for urban areas.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Dynamic 3d gaussian fields for urban areas

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.190261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.190261Z digest=sha256:c6be442cb11f4a48afb3f109e476b030b92901b744564a63a7f5aa684172ef3d

Observation 70eb2b3b-63eb-4257-a5a6-a8e64cbdd79a · outbound

This paper cites Ue4-nerf:neural radiance field for real-time rendering of large-scale scene.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Ue4-nerf:neural radiance field for real-time rendering of large-scale scene

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.253760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.253760Z digest=sha256:b6c79c60b579ae0934a8db740db15eb95e14cc387e5eddbcfc03ecca34e17010

Observation a5fa8581-1343-47b9-8e7c-a86977f9e80d · outbound

This paper cites StreetSurf: Extending Multi-view Implicit Surface Reconstruction to Street Views.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields StreetSurf: Extending Multi-view Implicit Surface Reconstruction to Street Views

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.307473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.307473Z digest=sha256:646d19bf07a33926b53579f8b9c9cdb29403c9751f48995780f1f497223ab111

Observation 36a3ba48-3582-43fb-8f84-603d4d76c0a2 · outbound

This paper cites Visual program- ming: Compositional visual reasoning without training.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Visual program- ming: Compositional visual reasoning without training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.378479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.378479Z digest=sha256:9434ba27f923703577a5e0ddb347c5b3d9b71e1547457419527ce6f2d3f6a1d3

Observation 15cf77ae-37cc-48a1-9156-f7661f6bd688 · outbound

This paper cites Pigeon: Predicting image geolocations.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Pigeon: Predicting image geolocations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.437098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.437098Z digest=sha256:7f51727a2b162d9cf85b00e2ffcd33c498458ce9e583723eedee206fa40237d0

Observation a4e7e80a-1a68-44a6-a34d-90d8b1bbb3b8 · outbound

This paper cites Dragon: Drone and ground gaussian splatting for 3d building reconstruction.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Dragon: Drone and ground gaussian splatting for 3d building reconstruction

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.490242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.490242Z digest=sha256:ff16587f8a0cda66b72c09cd11eebc1e105c6e022591324efaae2adfa81dad9a

Observation eb5e3a97-aac3-425d-a4cc-cc6dec241d00 · outbound

This paper cites 3d-llm: Inject- ing the 3d world into large language models.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 3d-llm: Inject- ing the 3d world into large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.533958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.533958Z digest=sha256:aea3bdeffa2c1064dff6d94afe1d9e47f8123e5202ed375c8e81a705d8fd7a68

Observation 3edb0cda-0e2e-45c5-b5c1-bc48b5709486 · outbound

This paper cites RSGPT: A Remote Sensing Vision Language Model and Benchmark.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.574175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.574175Z digest=sha256:40002c060a445209a0bbf54ce463576769645a27bcf558dff088066db3534c0f

Observation 57381278-8928-4fa2-bd2c-0c0df261d55d · outbound

This paper cites TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.614807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.614807Z digest=sha256:42a2a0501dc67b3031e2abde0026960f1ddeddde12581176111e6ff0b6c976a1

Observation cacaa869-fdd5-498b-aa6c-f7c052d768b2 · outbound

This paper cites GraspSplats: Efficient Manipulation with 3D Feature Splatting.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields GraspSplats: Efficient Manipulation with 3D Feature Splatting

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.660608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.660608Z digest=sha256:08e125ccff9068fc27c712fcdd03d9f526a5ca0518db90f73f070af8788d6add

Observation da349638-f488-43c6-9293-b7e9a72dbcaa · outbound

This paper cites Fastlgs: Speeding up lan- guage embedded gaussians with feature grid mapping.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Fastlgs: Speeding up lan- guage embedded gaussians with feature grid mapping

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:50:00.199756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:43.729855Z digest=sha256:7e65a57f285448251e311d964abc9c514f26cb6417c103dd5a17a4c00fe8271f

Observation 4cc0bab2-8de2-4525-a191-6972b1f1a3ef · outbound

This paper cites 3d gaussian splatting for real-time radiance field rendering.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 3d gaussian splatting for real-time radiance field rendering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:50:00.069090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:43.778265Z digest=sha256:182ac77d375b1d7fc4fbe03a338c2872bb5c0306597a7c730b318861d41caac5

Observation 5079a02f-46e7-4f71-96b1-0ae8264c004f · outbound

This paper cites A hierarchical 3d gaussian representation for real-time ren- dering of very large datasets.ACM Transactions on Graphics (TOG), 43(4), 2024.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields A hierarchical 3d gaussian representation for real-time ren- dering of very large datasets.ACM Transactions on Graphics (TOG), 43(4), 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.894873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:43.817893Z digest=sha256:b2bd4a53bafb4b7da11284512ad090f0ebd173bc65f1e8c4312b78d2c18c6b25

Observation f03dc12f-8688-4d4f-b78a-3b4d89abee4c · outbound

This paper cites LERF: Language embed- ded radiance fields.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields LERF: Language embed- ded radiance fields

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.718873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:43.885600Z digest=sha256:459e5f88fd84e51f9706bd2c12aebe48590a19714293aebc4612fcafaaf79701

Observation 4d690336-278a-482b-b22a-6b0373ac9ea7 · outbound

This paper cites Lobell, and Ste- fano Ermon.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Lobell, and Ste- fano Ermon

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.590614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:43.952849Z digest=sha256:523961a5bf61be1d73fd6c35ef1ec2528ff09abbf91dace8459fe9105eaa1e0f

Observation 77bafc28-a913-4c26-8aad-80d6f1db3910 · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.476659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:44.000320Z digest=sha256:ac06fe51216920dc834301d7e7803c79f2ce33091b8662f9c7e280624415173b

Observation 512e23fd-0770-4804-8f56-d514e6f22f0c · outbound

This paper cites Segment any- thing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Segment any- thing

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.303634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:44.042944Z digest=sha256:a409c62a9ac64280e204adc00835fc5684689b91bbc78415925111c949ea31d3

Observation 383964ea-5fb3-40a4-b019-7b71dbe8a3ed · outbound

This paper cites SatCLIP: Global, General-Purpose Location Embeddings with Satellite Imagery.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields SatCLIP: Global, General-Purpose Location Embeddings with Satellite Imagery

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:44.101737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:44.101737Z digest=sha256:a1796df6a249417e5a21e54d0598dd6a26514468b08991f56888163a0b7446c5

Observation db4b74fd-35ac-4f0d-96c4-308024254d38 · outbound

This paper cites Decomposing nerf for editing via feature field dis- tillation.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Decomposing nerf for editing via feature field dis- tillation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.166663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:44.136124Z digest=sha256:1c51703aaae2297305d759c8062bce4a6fdfef495e64896f5edcc20dddfccb22

Observation 4594861b-9cfa-4255-9ca0-32b88981eb59 · outbound

This paper cites Text2pos: Text-to-point-cloud cross-modal localiza- tion.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Text2pos: Text-to-point-cloud cross-modal localiza- tion

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.064991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:44.170793Z digest=sha256:aa260fb717e28b17048f60b758d0adfde7e0d1565c3775008d09681fb7ba3284

Observation 8eb4957f-fbb9-41b1-8285-94555bb82c3d · outbound

This paper cites Geochat: Grounded large vision-language model for remote sensing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Geochat: Grounded large vision-language model for remote sensing

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.903803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:44.222774Z digest=sha256:3782287401129b2ac493398ba125e6023b25b21097210dad6cd912ff42bd5552

Observation 7b1fd120-5535-440d-8b72-83d94a97fa55 · outbound

This paper cites NeRF-XL: Scaling nerfs with multiple GPUs.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields NeRF-XL: Scaling nerfs with multiple GPUs

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.773105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:44.267005Z digest=sha256:7eac1884dd2565e291046a48d333b68c9c854de268e90b782fcbdef14eae3589

Observation 51b0f3b3-7f83-4852-bc1d-694b87713b73 · outbound

This paper cites Matrixcity: A large-scale city dataset for city-scale neural rendering and beyond.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Matrixcity: A large-scale city dataset for city-scale neural rendering and beyond

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.613797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:44.309914Z digest=sha256:51203bbb285554da16e57b93405b3a22ccfda4cbfb7a65c8f76492ed7b3c305d

Observation a5a46439-664f-40d8-a915-3f451efd8976 · outbound

This paper cites Vastgaussian: Vast 3d gaus- sians for large scene reconstruction.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Vastgaussian: Vast 3d gaus- sians for large scene reconstruction

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.486658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:44.345633Z digest=sha256:7acdca1ef4a2128f16906fbf930d95af7c3d9afda7a58f252b62d2bf4127d5c7

Observation 11471b57-9a6a-442e-adb6-c54a6d7a3237 · outbound

This paper cites Capturing, reconstructing, and simulating: the urbanscene3d dataset.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Capturing, reconstructing, and simulating: the urbanscene3d dataset

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.345746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:44.377705Z digest=sha256:bd28d7ae6e87349bd0dcde247ddf0d4f473d6b11d0fb6ac5c50f229c98f17ead

Observation 2ffa59ba-1675-421a-89c6-c32d5bd75650 · outbound

This paper cites Re- moteclip: A vision language foundation model for remote sensing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Re- moteclip: A vision language foundation model for remote sensing

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.217038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:44.441781Z digest=sha256:6de29b0d2823cb68958e8285844515facdb2ea9cc9cb7b5a99c4d53e008b37a1

Observation 321362df-eaa4-42f8-ad62-66f8dfe04287 · outbound

This paper cites Visual instruction tuning.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Visual instruction tuning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.049953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:44.492129Z digest=sha256:b71906242868d014c4900e2eb340c33fa96cfa5ab839046f03f1654b8e141f84

Observation 7da5c1eb-1eb3-48c3-abb9-73bcdc066f97 · outbound

This paper cites Citygaus- sian: Real-time high-quality large-scale scene rendering with gaussians.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Citygaus- sian: Real-time high-quality large-scale scene rendering with gaussians

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:57.943733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:44.536960Z digest=sha256:f9437c28d892c500caa9c33365e1136ce65459ba9153dfa6f851f711f20d0eba

Observation c5b516e5-64a8-4bbc-9865-6fbfc9fb13b8 · outbound

This paper cites Citygaussianv2: Efficient and geometri- cally accurate reconstruction for large-scale scenes, 2024.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Citygaussianv2: Efficient and geometri- cally accurate reconstruction for large-scale scenes, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:57.732635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:44.579863Z digest=sha256:081c422dd1b8dc37df00fadbcafe40f10186ea4f90d8d4ccbaf07309baa6af94

Observation 64464ef6-4f58-491e-a64e-c1aea177f006 · outbound

This paper cites Citygaussianv2: Efficient and geometri- cally accurate reconstruction for large-scale scenes.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Citygaussianv2: Efficient and geometri- cally accurate reconstruction for large-scale scenes

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:57.486653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:44.633032Z digest=sha256:77f5e68ddb1078685fba1c71808f0fee043049a2e1f761c4c99a87a046e4fcc7

Observation a39d028d-c282-40a8-841b-91b5017f9523 · outbound

This paper cites Chameleon: Plug-and-play compositional reasoning with large language models.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Chameleon: Plug-and-play compositional reasoning with large language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:57.326550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:44.673401Z digest=sha256:39486e049939304330eb5ab250b6638e886bfdb362ee53128e8f31ad7dc2c9d7

Observation 411a2e79-1115-4057-b457-d2753476d1ce · outbound

This paper cites Exploring models and data for remote sensing im- age caption generation.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Exploring models and data for remote sensing im- age caption generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:57.153451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:44.722564Z digest=sha256:433f903f1cb66428108290e73646ac820f576dbc9422e4326cb46860b6129eb9

Observation 42f8f1c8-c23b-4d8e-bec0-fae4bce9c14e · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:44.779870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:44.779870Z digest=sha256:210a4a63d2af877260cd3fc2633de1c8dd0fea30c6e085ae1003c0fadc280d4b

Observation 1c23294a-ff1c-43ac-9251-0d118650d631 · outbound

This paper cites A multiscale grouping transformer with clip latents for re- mote sensing image captioning.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields A multiscale grouping transformer with clip latents for re- mote sensing image captioning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.999396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:44.839366Z digest=sha256:4691aae2fe825aaf98c2de345bbc3154386b27a0b7fefaf8cf73e443f2c89deb

Observation 8692614f-bbd4-425f-b298-23576a5f2c93 · outbound

This paper cites Llama 3.2 connect 2024: Vision on the edge and mo- bile devices.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Llama 3.2 connect 2024: Vision on the edge and mo- bile devices

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.865133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:44.886549Z digest=sha256:cb8b93119aef4de3c8f080b2fd4a94f2ca0b5d057821fc2294658be6b869efcd

Observation 968bc7a2-0782-4099-b787-0fc6d70410f8 · outbound

This paper cites Srinivasan, Matthew Tancik, Jonathan T.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Srinivasan, Matthew Tancik, Jonathan T

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.632184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:44.962424Z digest=sha256:15961173d4f12b09d6a14cdcdde1f3fe09891651d5665da2db882d0286644201

Observation dadfd791-b5fc-4459-a5cf-e26c4aa29d90 · outbound

This paper cites Cityrefer: Geography-aware 3d visual grounding dataset on city-scale point cloud data.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Cityrefer: Geography-aware 3d visual grounding dataset on city-scale point cloud data

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.486953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:45.026632Z digest=sha256:2ea221dc1bf05f5d5f721b3605a3807bf4f4d563176850490b3457dea1be98a3

Observation c84971c3-8ab1-4f8c-98e6-f143ba7b7feb · outbound

This paper cites Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.346668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:45.078369Z digest=sha256:382b4968e8c28322bc6674505f98572c6b6555af2dd171ea312ceb0235c77f89

Observation a6f53a50-2fbb-4fa8-9bb5-82087eb5ffff · outbound

This paper cites Hello gpt-4o.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Hello gpt-4o

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.242315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:45.171325Z digest=sha256:583335a82bc592f106aa94e8ce27ed5db099f549361c8773253a88e9a8ccd512

Observation a7d2eac8-8edb-49c0-9d47-9a7287cd4e45 · outbound

This paper cites Vhm: Versatile and honest vision lan- guage model for remote sensing image analysis.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Vhm: Versatile and honest vision lan- guage model for remote sensing image analysis

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.086960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:45.242656Z digest=sha256:a4942c5e25c279b45993e5b7ea7a6fd2e9c8139a3fccc3cf06b37e99ccf2f016

Observation ede4e48d-4a1f-44e3-8d07-8c976f4a375d · outbound

This paper cites Langsplat: 3d language gaussian splatting.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Langsplat: 3d language gaussian splatting

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.894524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:45.317392Z digest=sha256:8804f88de7bc303fdb51b448b59241d2b8a0cf3fd2b82f476a0b52762dc24df9

Observation 597031c2-4287-4768-9ec9-9c81916f7abe · outbound

This paper cites Deep semantic understanding of high resolution remote sensing image.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Deep semantic understanding of high resolution remote sensing image

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.736156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:45.454471Z digest=sha256:192fef0c555dbc736f47c654ee198939960e18db97d193441aca5380e2a9e19f

Observation c47a1362-4105-4afd-b80a-93a9f0db17c5 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Learn- ing transferable visual models from natural language super- vision

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.596461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:45.537922Z digest=sha256:729b89240d0d0d8bd368229d3aa5e5c3ea93d39bf5971a8fffdaed03760d2cec

Observation 12ebebed-4dce-4e81-9e20-ce61af8ba3db · outbound

This paper cites Derf: Decom- posed radiance fields.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Derf: Decom- posed radiance fields

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.445680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:45.650119Z digest=sha256:e5b2cb928b96dcec1ce56c1e74cb9f78c34a9c566dfa921e0e055022bd04e2e0

Observation 72561cb0-d64b-4d2e-9e0d-559e86cabb15 · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Toolformer: Language models can teach themselves to use tools

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.249562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:45.703626Z digest=sha256:2e74aa177fe6326175820b0b8cf77e7c987513a0439e4f560f5aa895bd456997

Observation c64e465f-305b-429a-8906-0612a0ab092e · outbound

This paper cites CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:45.758919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:45.758919Z digest=sha256:9dc1757ca3c6d223042c76f9e088fe702f9db7715a23eb20783a336da7072435

Observation 1cca6068-bdbe-4e95-b372-49d2d1cc4e29 · outbound

This paper cites Language embedded 3d gaussians for open- vocabulary scene understanding.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Language embedded 3d gaussians for open- vocabulary scene understanding

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.086891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:45.808762Z digest=sha256:3230765d4e85a186a21a0bb1a1f663dcc72353df8f3eb73e16037d3c1716002b

Observation 753f4515-7c61-4625-8f72-df9435fd83cc · outbound

This paper cites Real-time view synthesis for large scenes with millions of square meters.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Real-time view synthesis for large scenes with millions of square meters

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.947992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:45.867334Z digest=sha256:d5ced86ea159cd42d99ec99b28cfe06c3c3974a1feab59ad6527a53e434902fa

Observation 86280683-e13f-4125-a85d-f36cb77c5652 · outbound

This paper cites City-on-web: Real-time neural rendering of large- scale scenes on the web.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields City-on-web: Real-time neural rendering of large- scale scenes on the web

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.815656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:45.905136Z digest=sha256:13d959502d11ad1cdf65e85189b1fee9e9f3f8dfb369e122ebbb6960bf37d721

Observation b9c0206e-f8e7-4604-82a5-16898e815aeb · outbound

This paper cites De- composing 3d scenes into objects via unsupervised volume segmentation.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields De- composing 3d scenes into objects via unsupervised volume segmentation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.633027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:46.023299Z digest=sha256:643f1508e708066b187d9ed9735913803a1d792fea437b7082fb98062c89bc9e

Observation 700651a4-cb5f-4101-b462-ca96833df9cb · outbound

This paper cites Modular visual question answering via code generation.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Modular visual question answering via code generation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.454110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:46.069829Z digest=sha256:2d3a5365c62886fe6264e1830a8636a393681b98a895da8c2ce04943c16890f8

Observation dc431695-f1e0-4f6f-a2d7-2bbbf9863b43 · outbound

This paper cites 3d ques- tion answering for city scene understanding.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 3d ques- tion answering for city scene understanding

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.292203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:46.122863Z digest=sha256:c9af10c14360e0b06ceace4be9579ce0574a132c9167b5f2cf2b0fedc9b705dc

Observation a116152e-ee2d-47bc-a708-0b273c10d0b5 · outbound

This paper cites Visual grounding in remote sensing images.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Visual grounding in remote sensing images

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.192127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:46.179689Z digest=sha256:86514e09850e10b03c0cbbee43f0f2280e961495321b6581b3cb56c3f79e5c27

Observation 89523449-652f-4b7d-93aa-f2de235e1a3b · outbound

This paper cites ViperGPT: Visual inference via python execution for reasoning.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields ViperGPT: Visual inference via python execution for reasoning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.059622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:46.235926Z digest=sha256:e8bd70541cc4533ea4da8b6a02f2b8317f0330805ba3e65a75df7458b9fa59e0

Observation d3fd2ac5-28eb-43fd-947f-dc05078d97f3 · outbound

This paper cites Srinivasan, Jonathan T.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Srinivasan, Jonathan T

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:53.881502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:46.293702Z digest=sha256:6e4573aad2b7cdf877e2e6355b044b8a7b8cbdb53df17da805f86a537d7585e4

Observation 57e570d9-5479-4183-9948-0f315a385603 · outbound

This paper cites Crs-diff: Controllable generative re- mote sensing foundation model.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Crs-diff: Controllable generative re- mote sensing foundation model

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:53.729228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:46.366132Z digest=sha256:7c886ea63e3b060540bf0f87f104fb28610d6c7fbbe5766adeefd06724fccf49

Observation e963ca73-2e31-42fd-b422-939edfc5e32f · outbound

This paper cites MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:46.416409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:46.416409Z digest=sha256:f1cf82e0a5f8a2ccdb35fd386182a40d554b2fe23da77ee83c2e48cf607f9f84

Observation 43ac9ec0-b7ff-427a-a3de-96a9fb84d01e · outbound

This paper cites Mega-nerf: Scalable construction of large-scale nerfs for virtual fly-throughs.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Mega-nerf: Scalable construction of large-scale nerfs for virtual fly-throughs

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:53.611663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:46.462758Z digest=sha256:9e9e82d3c25b1cf8f1ff359b7c2d16b131c6824d0d9ac201696ed2aa8f68a589

Observation 66328b1d-9627-491e-b845-e8701b4ae3d7 · outbound

This paper cites Geoclip: Clip-inspired alignment between locations and im- ages for effective worldwide geo-localization.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Geoclip: Clip-inspired alignment between locations and im- ages for effective worldwide geo-localization

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:53.524417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:46.542012Z digest=sha256:96d14b67856cd28c0b72165f0703bff674a5827e179dc450bed851476b64b058

Observation 08dcae4e-331b-4fde-a88e-c931cd680623 · outbound

This paper cites Skyscript: A large and semanti- cally diverse vision-language dataset for remote sensing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Skyscript: A large and semanti- cally diverse vision-language dataset for remote sensing

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:53.281261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:46.608758Z digest=sha256:23f741a65270a82ec9c15f14adad41bfcabd7b2be31dbc82223aa0164905a870

Observation 96d58b29-89ee-43af-b61c-3fb6d54e869a · outbound

This paper cites Text2loc: 3d point cloud localization from natural language.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Text2loc: 3d point cloud localization from natural language

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:52.994039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:46.661207Z digest=sha256:3598e78e2c1063399426cc4647497fc4be1d61252f6f4fb9c7f9e0b9e3f019b4

Observation e6fdab54-f947-4033-8866-5f2f7e532a5c · outbound

This paper cites Bungeenerf: Progressive neural radiance field for extreme multi-scale scene rendering.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Bungeenerf: Progressive neural radiance field for extreme multi-scale scene rendering

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:52.782372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:46.717819Z digest=sha256:270f54349c9d58445f791ce1153ca3b3fed86c1fc8ad01e41ff052fa2524a63b

Observation 74ea19dd-d55c-4610-96da-75d4db71c88a · outbound

This paper cites Citydreamer: Compositional generative model of unbounded 3D cities.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Citydreamer: Compositional generative model of unbounded 3D cities

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:52.490533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:46.764874Z digest=sha256:b542dda4500e1cb3c5ce64efedb96f8a450cb93e30383595a214b299a16cd1cc

Observation 7cb1dfe4-9caa-4660-8e0f-322ee862b86d · outbound

This paper cites Generative Gaussian Splatting for Unbounded 3D City Generation.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Generative Gaussian Splatting for Unbounded 3D City Generation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:46.818968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:46.818968Z digest=sha256:0eb7459f5ade7c2beda9051159c177bda9aa9aa87466141a34dc71a90fb33768

Observation 8ad4195c-3cd5-4c51-9efb-b631cd4f3f37 · outbound

This paper cites Grid-guided neural radiance fields for large urban scenes.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Grid-guided neural radiance fields for large urban scenes

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:52.184167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:46.895366Z digest=sha256:a732984200252e31f08ec52a7646e41fc7f24f6bbdf9efc349f48acaceca58c5

Observation b760559c-52cb-4306-a908-3e67e5e05eeb · outbound

This paper cites Pointllm: Empowering large language models to understand point clouds.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Pointllm: Empowering large language models to understand point clouds

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:51.883576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:46.938093Z digest=sha256:2673e3ab289afb765b5d89ddafd6f4fccbee331b41ce5307c99fe64b9b41bb15

Observation e3dd9d04-2419-4c03-a6a6-a96cfc73e4c1 · outbound

This paper cites Addressclip: Empowering vision-language models for city-wide image address localization.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Addressclip: Empowering vision-language models for city-wide image address localization

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:51.674509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:47.105618Z digest=sha256:4aeefc6ce0b1feb0109867e67a4c0671e0fc9e2d1fc9d222c5f1ce73787c1dc9

Observation c5b83c66-8f0f-4dff-983b-606d80e0b1ee · outbound

This paper cites Unisim: A neural closed-loop sensor simulator.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Unisim: A neural closed-loop sensor simulator

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:51.420687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:47.217718Z digest=sha256:a0be1620083879365fc3d89a8b2795acc05215f7c12b9caaac3f55de4313f1e2

Observation 8a8f6348-3a1f-4b59-8b62-811d51aae34f · outbound

This paper cites Scannet++: A high-fidelity dataset of 3d indoor scenes.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Scannet++: A high-fidelity dataset of 3d indoor scenes

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:51.161037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:47.345334Z digest=sha256:c55d1b6c068b8ac6ae6742ff98b31d49932699872b825d40989bb141ca20fa71

Observation 928e6764-3973-4ad1-bf7b-d15f99d7e2e9 · outbound

This paper cites Dogs: Distributed-oriented gaus- sian splatting for large-scale 3d reconstruction via gaussian consensus.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Dogs: Distributed-oriented gaus- sian splatting for large-scale 3d reconstruction via gaussian consensus

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:50.891902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:47.442586Z digest=sha256:f43df44b3d10b36abeacaf2ed8594283aef3f71b75a6103e3660737cab49d8a1

Observation c06fb373-00e8-4647-aca5-bfeb0b7331a2 · outbound

This paper cites PreSight: Enhancing Autonomous Vehicle Perception with City-Scale NeRF Priors.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields PreSight: Enhancing Autonomous Vehicle Perception with City-Scale NeRF Priors

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:49:48.524762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:47.509275Z digest=sha256:4b475b16bf13c76fba67e3e5f6d20be0bb9101ac8af20d761a2a131e2097f79f

Observation 07bb5d59-1330-4260-90f3-f1ce37af3872 · outbound

This paper cites Exploring a fine-grained multiscale method for cross-modal remote sensing image re- trieval.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Exploring a fine-grained multiscale method for cross-modal remote sensing image re- trieval

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:50.727057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:47.551158Z digest=sha256:dd7425bab6d954be49e0cda8a5c4b5d4d844d5ba3afc0126fc3a7a0b748117bf

Observation c8b06b90-3f35-40f9-bfbd-76a741394e33 · outbound

This paper cites Garfield++: Reinforced gaussian ra- diance fields for large-scale 3d scene reconstruction, 2024.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Garfield++: Reinforced gaussian ra- diance fields for large-scale 3d scene reconstruction, 2024

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:50.489234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:47.607078Z digest=sha256:98d2d9e5da63cc923033a44a9abc4ca09475c5313016a232e6bd6437d85da960

Observation 1444202d-ec7b-4a38-9b77-da8efe02732a · outbound

This paper cites 3DitScene: Editing any scene via language-guided disentan- gled gaussian splatting.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 3DitScene: Editing any scene via language-guided disentan- gled gaussian splatting

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:50.181200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:47.663922Z digest=sha256:f050be88e9b40f01020f1c9526bcaaa9e1028edab92581ff442e08a40c70aa73

Observation 5d435d0e-5c3e-4e41-bda2-42231a60e7fa · outbound

This paper cites Earthgpt: A universal multi-modal large lan- guage model for multi-sensor image comprehension in re- mote sensing domain.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Earthgpt: A universal multi-modal large lan- guage model for multi-sensor image comprehension in re- mote sensing domain

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:49.989182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:47.719996Z digest=sha256:099d532a3db9e7f103d903dbac8cf6ccdea59c8dd55cdeba77f0783bddbf612e

Observation cb0f6bdb-ae78-497e-8085-d4022ea2b322 · outbound

This paper cites EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:47.792628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:47.792628Z digest=sha256:311580c6f504f18798f554eae75b4233413d3a39b5a258499672a858a7835cf1

Observation a6fde006-503c-4e13-a313-7da9fb1a1243 · outbound

This paper cites Efficient Large-scale Scene Representation with a Hybrid of High-resolution Grid and Plane Features.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Efficient Large-scale Scene Representation with a Hybrid of High-resolution Grid and Plane Features

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:49:48.374875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:47.847955Z digest=sha256:ea11f3ab60a9e5e059c4cade1d7455c9c724ef1df34b8e7848737f6d0d6ac05d

Observation a7ef6112-0d56-40b4-8e46-add371eadc78 · outbound

This paper cites Aerial lifting: Neural urban semantic and building instance lifting from aerial imagery.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Aerial lifting: Neural urban semantic and building instance lifting from aerial imagery

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:49.770435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:47.900648Z digest=sha256:f27cd14a213b3bd313494366636bd75bfabedc316632e8a554d174b15aca53e4

Observation de21f140-582b-4c75-9827-6bbac2003409 · outbound

This paper cites Rs5m and georsclip: A large-scale vision- language dataset and a large vision-language model for remote sensing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Rs5m and georsclip: A large-scale vision- language dataset and a large vision-language model for remote sensing

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:49.525950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:47.958787Z digest=sha256:06be5e1c9a45dc413b53c9ea98ed5f6cf89919238f4440b1b2f0854075721250

Observation 861b8f96-6aef-48a5-9807-82467594e760 · outbound

This paper cites Mutual Attention Inception Network for Remote Sensing Visual Question Answering.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Mutual Attention Inception Network for Remote Sensing Visual Question Answering

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:49.261423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:48.037601Z digest=sha256:b8ecdbc0b8f5b7caf53ec417870a405be4c52ed62db4aafa6ee6a8513d7e335a

Observation e9afdf7b-ffbb-44f9-8cac-136746af8ab1 · outbound

This paper cites Drivinggaussian: Composite gaussian splatting for surrounding dynamic au- tonomous driving scenes.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Drivinggaussian: Composite gaussian splatting for surrounding dynamic au- tonomous driving scenes

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:49.088711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:48.120675Z digest=sha256:03e0b7a7621db2a8745fc1192bb88faab1091ec645773e25b7d33cb5c289ee17

Observation 067fb573-df17-4484-9755-f991eb3f2950 · outbound

This paper cites Towards vision- language geo-foundation models: A survey.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Towards vision- language geo-foundation models: A survey

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:48.162222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:48.162222Z digest=sha256:cebc7a6b997845451e93010edbfbf98f3741fbd5982cb9e00950543f5f35529b

Observation 665b8c4d-6538-4db8-9d99-a80c8c9f3ed5 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:48.941651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:48.227457Z digest=sha256:7694daa71ca5cf65ed7dba23f7fd3cd550c3b2b6b750cc07049603610e3993c0

Observation 0eddabe3-4fb5-4160-a229-5338d77d636e · outbound

This paper cites 'yes' if {ANSWER1} < {ANSWER2} else 'no'.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 'yes' if {ANSWER1} < {ANSWER2} else 'no'

Reference 2023

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T21:49:48.805064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T21:49:48.230001Z digest=sha256:074835379478c9e50f234dccb32bcc5761770ec64ff02c4050956f978d2686ff

Pith citing papers

No inbound Pith citation observations are available.