Pith. sign in

Paper Citation Record · LEDGER

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields

As of 21 August 2026, this Paper Citation Record lists 96 of 96 outbound references and 0 inbound Pith citation observations for arXiv:2506.23352.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23352 v1

Coverage vector

measured 96 of 96 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:49:48.230001Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

96 of 96 outbound references displayed

  • verified exact2
  • verified fuzzy64
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9a2a456f-48b1-461b-9bab-b5febf5309c0 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Scanqa: 3d question answering for spatial scene understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.054070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.054070Z digest=sha256:7674f7fbbb64de95ec441e5e1440119e3d86bb836df592941960811171c26cc9

Observation 7bba69c6-6d78-4fd5-89d6-9cdd55d2ada1 · outbound

This paper cites Qwen2.5-VL Technical Report.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.094181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.094181Z digest=sha256:b2805472f59379af1d4ba61c720354acbb8aaa351b59831ef87b3d1e7a1a6a81

Observation 3922716b-968a-46df-9b06-41652b632605 · outbound

This paper cites Henriques, Andrew Zisserman, and Andrea Vedaldi.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Henriques, Andrew Zisserman, and Andrea Vedaldi

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.194838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.194838Z digest=sha256:30175dc86e39180694edd7ac4130db73c6eea470435b0332eee01e3c2a7413d2

Observation 05a9f140-72af-403e-8111-ef8f7d6c539b · outbound

This paper cites Blumer, Qingx- uan Chen, and Francis Engelmann.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Blumer, Qingx- uan Chen, and Francis Engelmann

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.256991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.256991Z digest=sha256:9373c1136228b99c849c1971dbbc3e6791813e294a28a18b91f22cadcc657d67

Observation b093582a-470c-4d28-b589-ce745a4c8e3c · outbound

This paper cites A persistent spatial semantic representation for high-level natural language instruction execution.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields A persistent spatial semantic representation for high-level natural language instruction execution

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.392137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.392137Z digest=sha256:9b4da3dc11ffa62cc387046398e553cb82a610502c57a31c2d14eed4b6504032

Observation 71ca8a59-5ff8-47a7-9292-8244e40f89d9 · outbound

This paper cites Prompt-rsvqa: Prompt- ing visual context to a language model for remote sensing visual question answering.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Prompt-rsvqa: Prompt- ing visual context to a language model for remote sensing visual question answering

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.499765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.499765Z digest=sha256:ec5d33b6566832c7dfd38032849f9a92c32055416b6ef8c7532af1abbeff69d7

Observation dee5502b-394d-4b8a-9d94-53bcfd086c67 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natu- ral language.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Scanrefer: 3d object localization in rgb-d scans using natu- ral language

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.634508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.634508Z digest=sha256:aba0e5423c4a9fe2364d7138a59af7ab925c165b4d1efd8302a3cb0b86b96808

Observation 2d18a5ce-27c7-4499-9d5a-6992720f8660 · outbound

This paper cites Panoptic vision-language feature fields.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Panoptic vision-language feature fields

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.745471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.745471Z digest=sha256:70fb192ca940931cfe59136b5b144f9bbcccb320d8b781fad822bb4bd8adabff

Observation e45035e5-1c0c-469c-a16e-8289272ccb08 · outbound

This paper cites Stylecity: Large-scale 3d urban scenes stylization.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Stylecity: Large-scale 3d urban scenes stylization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.839057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.839057Z digest=sha256:5eefa7e83ab504515059d62d66c2dc4b585cd27575288082df222ee924f7b21d

Observation d71297e6-2235-430c-96e6-bf273376ff1e · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:42.929873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:42.929873Z digest=sha256:21292b316e9914faa9dca19f7471da7e0ef55c19b0604d25da8e43e45c1ba453

Observation 00fe7b97-3e74-4516-ab3c-9da16c510e9c · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.012679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.012679Z digest=sha256:73c6e22eac6f26ecc47cadd3706386e481adfe267aaa6f34f3ed192a4e7914ae

Observation da1aa099-20a3-46e0-a396-41eb294e5e9e · outbound

This paper cites LayoutGPT: Compositional Visual Planning and Generation with Large Language Models.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields LayoutGPT: Compositional Visual Planning and Generation with Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.132312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.132312Z digest=sha256:90c334e1ffb86c4ab44c7bc99c2c1ac9393bb4adf423a1824d2b8b598ca72de5

Observation 1999571b-cd9a-407c-bc5f-ccfc1b13af7e · outbound

This paper cites Dynamic 3d gaussian fields for urban areas.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Dynamic 3d gaussian fields for urban areas

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.190261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.190261Z digest=sha256:6201d8cdede1eeaa41d0b734103bb68ff689d957b53197f2508b51b7379843cf

Observation 70eb2b3b-63eb-4257-a5a6-a8e64cbdd79a · outbound

This paper cites Ue4-nerf:neural radiance field for real-time rendering of large-scale scene.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Ue4-nerf:neural radiance field for real-time rendering of large-scale scene

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.253760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.253760Z digest=sha256:78029c0cbfc13e48114e8acf472a118c2099306ef9975acaff7ca4985f339c97

Observation a5fa8581-1343-47b9-8e7c-a86977f9e80d · outbound

This paper cites StreetSurf: Extending Multi-view Implicit Surface Reconstruction to Street Views.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields StreetSurf: Extending Multi-view Implicit Surface Reconstruction to Street Views

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.307473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.307473Z digest=sha256:26971ed7cd565f481400aee71c7b9bd0afd8ea2820b40e3b2b5de85af17c1399

Observation 36a3ba48-3582-43fb-8f84-603d4d76c0a2 · outbound

This paper cites Visual program- ming: Compositional visual reasoning without training.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Visual program- ming: Compositional visual reasoning without training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.378479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.378479Z digest=sha256:fdcbec6bb09d7a19991bfcbb2179171caa7e5f2615a1c8a4a49a5133df0bedc1

Observation 15cf77ae-37cc-48a1-9156-f7661f6bd688 · outbound

This paper cites Pigeon: Predicting image geolocations.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Pigeon: Predicting image geolocations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.437098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.437098Z digest=sha256:17351a1d337b8788bb944c0506f452bfa528f2ff85aa805463a02a6911933980

Observation a4e7e80a-1a68-44a6-a34d-90d8b1bbb3b8 · outbound

This paper cites Dragon: Drone and ground gaussian splatting for 3d building reconstruction.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Dragon: Drone and ground gaussian splatting for 3d building reconstruction

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.490242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.490242Z digest=sha256:c64e93b6bd870bc6d2ae0fe1208b035a17b9b4c12ec62fb8215a73078aeaca35

Observation eb5e3a97-aac3-425d-a4cc-cc6dec241d00 · outbound

This paper cites 3d-llm: Inject- ing the 3d world into large language models.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 3d-llm: Inject- ing the 3d world into large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.533958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.533958Z digest=sha256:f1a61c8e4496a3be12c927e667ca71e9b3021f35616921f58fec648f7007f883

Observation 3edb0cda-0e2e-45c5-b5c1-bc48b5709486 · outbound

This paper cites RSGPT: A Remote Sensing Vision Language Model and Benchmark.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.574175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.574175Z digest=sha256:b3696867924dbe45d7c071bfb1cabaca917978efb1ffe8b23359a6f9cea5e77e

Observation 57381278-8928-4fa2-bd2c-0c0df261d55d · outbound

This paper cites TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.614807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.614807Z digest=sha256:85dcbd27c0c0519cbaaa999529681af5970251f2ee5e3da57c828a6bd8a12a0f

Observation cacaa869-fdd5-498b-aa6c-f7c052d768b2 · outbound

This paper cites GraspSplats: Efficient Manipulation with 3D Feature Splatting.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields GraspSplats: Efficient Manipulation with 3D Feature Splatting

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:43.660608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:43.660608Z digest=sha256:c4018f7ccae5434f2a1e6676980ad73cb93daab4adf00dae1a78b36577b7f213

Observation da349638-f488-43c6-9293-b7e9a72dbcaa · outbound

This paper cites Fastlgs: Speeding up lan- guage embedded gaussians with feature grid mapping.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Fastlgs: Speeding up lan- guage embedded gaussians with feature grid mapping

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:50:00.199756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:43.729855Z digest=sha256:b46282cc0ae5fa6c578c21a258ee71fadc5c0e8821b8fd402185efaf5b97f37f

Observation 4cc0bab2-8de2-4525-a191-6972b1f1a3ef · outbound

This paper cites 3d gaussian splatting for real-time radiance field rendering.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 3d gaussian splatting for real-time radiance field rendering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:50:00.069090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:43.778265Z digest=sha256:152610d434b9303e5cf3fdb4db6d6a2b4beb7b55c7772a1fe9686fee630ae3a1

Observation 5079a02f-46e7-4f71-96b1-0ae8264c004f · outbound

This paper cites A hierarchical 3d gaussian representation for real-time ren- dering of very large datasets.ACM Transactions on Graphics (TOG), 43(4), 2024.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields A hierarchical 3d gaussian representation for real-time ren- dering of very large datasets.ACM Transactions on Graphics (TOG), 43(4), 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.894873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:43.817893Z digest=sha256:7c2842ffc51044c0060e9669402842a5ae236da4f21139c46440fb1d978b57da

Observation f03dc12f-8688-4d4f-b78a-3b4d89abee4c · outbound

This paper cites LERF: Language embed- ded radiance fields.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields LERF: Language embed- ded radiance fields

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.718873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:43.885600Z digest=sha256:984ea898ee2b848226b431545466a9bc3d8fe2f8d5b2a47f215298ade0828594

Observation 4d690336-278a-482b-b22a-6b0373ac9ea7 · outbound

This paper cites Lobell, and Ste- fano Ermon.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Lobell, and Ste- fano Ermon

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.590614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:43.952849Z digest=sha256:0a0352bf194af5dedcab5faa612bba7243ab9926030c8bd4cc3a8a410affd78b

Observation 77bafc28-a913-4c26-8aad-80d6f1db3910 · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.476659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:44.000320Z digest=sha256:1eb57abebd2b570f3a05d6438469d7d1e5ae26a42308f59ed416853ab4ce2fd7

Observation 512e23fd-0770-4804-8f56-d514e6f22f0c · outbound

This paper cites Segment any- thing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Segment any- thing

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.303634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:44.042944Z digest=sha256:bc89c75ad6358faed3460e4c1d7a965a63fa351a59f346dcd5d32a2f171393fd

Observation 383964ea-5fb3-40a4-b019-7b71dbe8a3ed · outbound

This paper cites SatCLIP: Global, General-Purpose Location Embeddings with Satellite Imagery.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields SatCLIP: Global, General-Purpose Location Embeddings with Satellite Imagery

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:44.101737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:44.101737Z digest=sha256:d36bed6f18ec8fac371c9fa905b0ad63a6933fd9b2a180debee97bc12e3b8c76

Observation db4b74fd-35ac-4f0d-96c4-308024254d38 · outbound

This paper cites Decomposing nerf for editing via feature field dis- tillation.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Decomposing nerf for editing via feature field dis- tillation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.166663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:44.136124Z digest=sha256:c3402dbaab38b3d23961fef3f26c870a9c3da2d24c1a18658e91de20171ab258

Observation 4594861b-9cfa-4255-9ca0-32b88981eb59 · outbound

This paper cites Text2pos: Text-to-point-cloud cross-modal localiza- tion.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Text2pos: Text-to-point-cloud cross-modal localiza- tion

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:59.064991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:44.170793Z digest=sha256:5f61eacb342d5e44baa4d05f838005b606ae266ac249967d035ee49bc4e229aa

Observation 8eb4957f-fbb9-41b1-8285-94555bb82c3d · outbound

This paper cites Geochat: Grounded large vision-language model for remote sensing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Geochat: Grounded large vision-language model for remote sensing

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.903803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:44.222774Z digest=sha256:a3f9e52829269eb7bac776ddc170e3ca5b14520c6656995c29f9a6a5914c95da

Observation 7b1fd120-5535-440d-8b72-83d94a97fa55 · outbound

This paper cites NeRF-XL: Scaling nerfs with multiple GPUs.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields NeRF-XL: Scaling nerfs with multiple GPUs

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.773105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:44.267005Z digest=sha256:4e0ec216c5cccbff774977ee4184fc46e6007955acc31fb3bbf29c4607fdd9b4

Observation 51b0f3b3-7f83-4852-bc1d-694b87713b73 · outbound

This paper cites Matrixcity: A large-scale city dataset for city-scale neural rendering and beyond.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Matrixcity: A large-scale city dataset for city-scale neural rendering and beyond

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.613797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:44.309914Z digest=sha256:0327e1e11b7bb635a791b21cdd44db69e38a12353e6bdee4b84c54b05ed3a915

Observation a5a46439-664f-40d8-a915-3f451efd8976 · outbound

This paper cites Vastgaussian: Vast 3d gaus- sians for large scene reconstruction.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Vastgaussian: Vast 3d gaus- sians for large scene reconstruction

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.486658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:44.345633Z digest=sha256:59bca94f46dee2e989c1b650fd56d2ae51ff3e581dfaaf23adeaa8b07104292f

Observation 11471b57-9a6a-442e-adb6-c54a6d7a3237 · outbound

This paper cites Capturing, reconstructing, and simulating: the urbanscene3d dataset.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Capturing, reconstructing, and simulating: the urbanscene3d dataset

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.345746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:44.377705Z digest=sha256:dadd9ab59b1e4a90ac87429fbb4725923abe7493c113fc7b1f61e12165dfd515

Observation 2ffa59ba-1675-421a-89c6-c32d5bd75650 · outbound

This paper cites Re- moteclip: A vision language foundation model for remote sensing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Re- moteclip: A vision language foundation model for remote sensing

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.217038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:44.441781Z digest=sha256:0863a7ca6be8ede8a2e1f97f6e5aaa20fbb92aa7670ad41b423460abc61cd165

Observation 321362df-eaa4-42f8-ad62-66f8dfe04287 · outbound

This paper cites Visual instruction tuning.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Visual instruction tuning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:58.049953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:44.492129Z digest=sha256:067597c8c2f390d6b716d018f5ad90b5b10595b9c0e3bc1889b2dde6ab9b9c41

Observation 7da5c1eb-1eb3-48c3-abb9-73bcdc066f97 · outbound

This paper cites Citygaus- sian: Real-time high-quality large-scale scene rendering with gaussians.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Citygaus- sian: Real-time high-quality large-scale scene rendering with gaussians

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:57.943733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:44.536960Z digest=sha256:c031700dbf34125462707c9a3a9c2c50718501dc25704210d1505d0da9326944

Observation c5b516e5-64a8-4bbc-9865-6fbfc9fb13b8 · outbound

This paper cites Citygaussianv2: Efficient and geometri- cally accurate reconstruction for large-scale scenes, 2024.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Citygaussianv2: Efficient and geometri- cally accurate reconstruction for large-scale scenes, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:57.732635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:44.579863Z digest=sha256:28ddfaaf70559d77f993a49642462720fb29a5a8d6d3141fb8c72e6521ff2fb3

Observation 64464ef6-4f58-491e-a64e-c1aea177f006 · outbound

This paper cites Citygaussianv2: Efficient and geometri- cally accurate reconstruction for large-scale scenes.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Citygaussianv2: Efficient and geometri- cally accurate reconstruction for large-scale scenes

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:57.486653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:44.633032Z digest=sha256:c84aea8761f12ab7e6ace89cf862ca7c709354af4590fb9057790d105999b633

Observation a39d028d-c282-40a8-841b-91b5017f9523 · outbound

This paper cites Chameleon: Plug-and-play compositional reasoning with large language models.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Chameleon: Plug-and-play compositional reasoning with large language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:57.326550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:44.673401Z digest=sha256:a778f7985efc144ec2d341e7d913a98efa71d5af73610bb99722807c91df9552

Observation 411a2e79-1115-4057-b457-d2753476d1ce · outbound

This paper cites Exploring models and data for remote sensing im- age caption generation.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Exploring models and data for remote sensing im- age caption generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:57.153451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:44.722564Z digest=sha256:3ab727978d9e319981a4c67fbb455dc1fb16bd2d1e75a044d58cf012000c1f35

Observation 42f8f1c8-c23b-4d8e-bec0-fae4bce9c14e · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:44.779870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:44.779870Z digest=sha256:052b357a7b7766e2d7d0850cc5448f4392ebe9cc7c7de047d8882a090b4e8d06

Observation 1c23294a-ff1c-43ac-9251-0d118650d631 · outbound

This paper cites A multiscale grouping transformer with clip latents for re- mote sensing image captioning.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields A multiscale grouping transformer with clip latents for re- mote sensing image captioning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.999396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:44.839366Z digest=sha256:22df37fa92e6e30b645460fe7e77759d9221f0c647d71910246ae58509ff6b2c

Observation 8692614f-bbd4-425f-b298-23576a5f2c93 · outbound

This paper cites Llama 3.2 connect 2024: Vision on the edge and mo- bile devices.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Llama 3.2 connect 2024: Vision on the edge and mo- bile devices

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.865133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:44.886549Z digest=sha256:92356c06235b483fa010d2430fce9ba392f201dd9fd05a75e92ec80507ceadbf

Observation 968bc7a2-0782-4099-b787-0fc6d70410f8 · outbound

This paper cites Srinivasan, Matthew Tancik, Jonathan T.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Srinivasan, Matthew Tancik, Jonathan T

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.632184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:44.962424Z digest=sha256:33ac79ac83a1de73a554d010859f1a9d9f4b820e98944a6f18381a91c8cdf37f

Observation dadfd791-b5fc-4459-a5cf-e26c4aa29d90 · outbound

This paper cites Cityrefer: Geography-aware 3d visual grounding dataset on city-scale point cloud data.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Cityrefer: Geography-aware 3d visual grounding dataset on city-scale point cloud data

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.486953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:45.026632Z digest=sha256:96b547c8f9eb4306b126161b0afeaac1ddecb64476d6b741a2f5d06b01083621

Observation c84971c3-8ab1-4f8c-98e6-f143ba7b7feb · outbound

This paper cites Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Lhrs-bot: Empowering remote sensing with vgi-enhanced large multimodal language model

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.346668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:45.078369Z digest=sha256:e98aed9943ab4e2caf6a89d33d6634907ca4d91a0f8ee4f4f11b643001911cf1

Observation a6f53a50-2fbb-4fa8-9bb5-82087eb5ffff · outbound

This paper cites Hello gpt-4o.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Hello gpt-4o

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.242315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:45.171325Z digest=sha256:66ec870bd0c682c7fa633e2a889da31c11e0aee5048f13902a80a14c91fe4769

Observation a7d2eac8-8edb-49c0-9d47-9a7287cd4e45 · outbound

This paper cites Vhm: Versatile and honest vision lan- guage model for remote sensing image analysis.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Vhm: Versatile and honest vision lan- guage model for remote sensing image analysis

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:56.086960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:45.242656Z digest=sha256:c3954bbb5fd8634c1126524c02a61516c92819dcb977da12baa48b06ed878661

Observation ede4e48d-4a1f-44e3-8d07-8c976f4a375d · outbound

This paper cites Langsplat: 3d language gaussian splatting.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Langsplat: 3d language gaussian splatting

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.894524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:45.317392Z digest=sha256:fcaac70b92fe81711c72514e069de9281e23c29f330bbf75922fb793f7db2b01

Observation 597031c2-4287-4768-9ec9-9c81916f7abe · outbound

This paper cites Deep semantic understanding of high resolution remote sensing image.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Deep semantic understanding of high resolution remote sensing image

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.736156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:45.454471Z digest=sha256:107b9b71b2d7cc22fb03f3362871d4670919b7f7d0b17d0e72014ca90f47506d

Observation c47a1362-4105-4afd-b80a-93a9f0db17c5 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Learn- ing transferable visual models from natural language super- vision

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.596461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:45.537922Z digest=sha256:e83a3afc13d12287b9f69992dc119b1360055378c466bde63ff949d27a456d65

Observation 12ebebed-4dce-4e81-9e20-ce61af8ba3db · outbound

This paper cites Derf: Decom- posed radiance fields.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Derf: Decom- posed radiance fields

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.445680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:45.650119Z digest=sha256:1a17542eaf68f67e06f0f89b59a9a223a4d753bbe66244fad9be2dca3586ab41

Observation 72561cb0-d64b-4d2e-9e0d-559e86cabb15 · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Toolformer: Language models can teach themselves to use tools

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.249562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:45.703626Z digest=sha256:24e2dc76674742c5142aafd7d5afa3fdf8e2c8b4789e07432907327af385f4bb

Observation c64e465f-305b-429a-8906-0612a0ab092e · outbound

This paper cites CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:45.758919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:45.758919Z digest=sha256:87c16d63d465a07e553e0de757cf77d99841cd4a89bec3b5ca156884e865b3a8

Observation 1cca6068-bdbe-4e95-b372-49d2d1cc4e29 · outbound

This paper cites Language embedded 3d gaussians for open- vocabulary scene understanding.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Language embedded 3d gaussians for open- vocabulary scene understanding

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:55.086891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:45.808762Z digest=sha256:a0058d00afc1f14905824fd0d18923ee87c6c09abc6faa945c2fd1d0ad30dd65

Observation 753f4515-7c61-4625-8f72-df9435fd83cc · outbound

This paper cites Real-time view synthesis for large scenes with millions of square meters.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Real-time view synthesis for large scenes with millions of square meters

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.947992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:45.867334Z digest=sha256:1dec9d387b8ccacb7e11385910fced7dc80f12728867a71f27189237d6f8aa2d

Observation 86280683-e13f-4125-a85d-f36cb77c5652 · outbound

This paper cites City-on-web: Real-time neural rendering of large- scale scenes on the web.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields City-on-web: Real-time neural rendering of large- scale scenes on the web

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.815656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:45.905136Z digest=sha256:cad0f6df3b6f9fe0f4722206ef078a0bec5358a93a70435424b370233816bacf

Observation b9c0206e-f8e7-4604-82a5-16898e815aeb · outbound

This paper cites De- composing 3d scenes into objects via unsupervised volume segmentation.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields De- composing 3d scenes into objects via unsupervised volume segmentation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.633027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:46.023299Z digest=sha256:81ea263ff85d9383adcf0788348d74cc49576be4787818fe31db9c4fb9ac818d

Observation 700651a4-cb5f-4101-b462-ca96833df9cb · outbound

This paper cites Modular visual question answering via code generation.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Modular visual question answering via code generation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.454110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:46.069829Z digest=sha256:ba27d505ebd30b1fa00e415e7a633fed68ded40e21ef40e0ef7ee43262d7a184

Observation dc431695-f1e0-4f6f-a2d7-2bbbf9863b43 · outbound

This paper cites 3d ques- tion answering for city scene understanding.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 3d ques- tion answering for city scene understanding

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.292203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:46.122863Z digest=sha256:1cac24a257620b65a7b7d7283475c43be0805f982178d8e7fa2948a02d66bd6a

Observation a116152e-ee2d-47bc-a708-0b273c10d0b5 · outbound

This paper cites Visual grounding in remote sensing images.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Visual grounding in remote sensing images

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.192127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:46.179689Z digest=sha256:dad52a2a6d5ec87346cca0a071cbcbd347b419b22472c66d9292f209e19794d9

Observation 89523449-652f-4b7d-93aa-f2de235e1a3b · outbound

This paper cites ViperGPT: Visual inference via python execution for reasoning.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields ViperGPT: Visual inference via python execution for reasoning

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:54.059622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:46.235926Z digest=sha256:ad9c686bacacb309ca565bf7c9bf587631c6d6e8564a8222acb6ff300d398dbf

Observation d3fd2ac5-28eb-43fd-947f-dc05078d97f3 · outbound

This paper cites Srinivasan, Jonathan T.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Srinivasan, Jonathan T

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:53.881502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:46.293702Z digest=sha256:c2d51b6a4487ca1093584a3ad0156564f942fd222f72fa9e606ef7683fdeb532

Observation 57e570d9-5479-4183-9948-0f315a385603 · outbound

This paper cites Crs-diff: Controllable generative re- mote sensing foundation model.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Crs-diff: Controllable generative re- mote sensing foundation model

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:53.729228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:46.366132Z digest=sha256:573fc137519fcf088ef3718ee436e669856e362b4eb284850a44854daff7c6f1

Observation e963ca73-2e31-42fd-b422-939edfc5e32f · outbound

This paper cites MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:46.416409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:46.416409Z digest=sha256:a1ecedcc43e24071f3b157b80dcb86b978b34a6c627715683d40376cd0838028

Observation 43ac9ec0-b7ff-427a-a3de-96a9fb84d01e · outbound

This paper cites Mega-nerf: Scalable construction of large-scale nerfs for virtual fly-throughs.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Mega-nerf: Scalable construction of large-scale nerfs for virtual fly-throughs

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:53.611663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:46.462758Z digest=sha256:b1bb7a802fa26043372402b3b12b73bb1de4a7bffcd14cb973998b13a74161e5

Observation 66328b1d-9627-491e-b845-e8701b4ae3d7 · outbound

This paper cites Geoclip: Clip-inspired alignment between locations and im- ages for effective worldwide geo-localization.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Geoclip: Clip-inspired alignment between locations and im- ages for effective worldwide geo-localization

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:53.524417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:46.542012Z digest=sha256:e0354017a16aac018d386c29c316f98e2b47e9b6246ba712da2952af02ae2e14

Observation 08dcae4e-331b-4fde-a88e-c931cd680623 · outbound

This paper cites Skyscript: A large and semanti- cally diverse vision-language dataset for remote sensing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Skyscript: A large and semanti- cally diverse vision-language dataset for remote sensing

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:53.281261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:46.608758Z digest=sha256:840b71ae116ca1ac282a2d2a0d533b842fab5f693747d62e9083904020a5494c

Observation 96d58b29-89ee-43af-b61c-3fb6d54e869a · outbound

This paper cites Text2loc: 3d point cloud localization from natural language.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Text2loc: 3d point cloud localization from natural language

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:52.994039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:46.661207Z digest=sha256:b053a5432c7af72431dc9c6a264463187a923087e16c40664b4fb8a4a39cffe7

Observation e6fdab54-f947-4033-8866-5f2f7e532a5c · outbound

This paper cites Bungeenerf: Progressive neural radiance field for extreme multi-scale scene rendering.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Bungeenerf: Progressive neural radiance field for extreme multi-scale scene rendering

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:52.782372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:46.717819Z digest=sha256:f2b22d705ff99746a94ac57b0be1469bc4f7ed60bdda862eaa985ea415fa8248

Observation 74ea19dd-d55c-4610-96da-75d4db71c88a · outbound

This paper cites Citydreamer: Compositional generative model of unbounded 3D cities.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Citydreamer: Compositional generative model of unbounded 3D cities

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:52.490533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:46.764874Z digest=sha256:8dddc965e0a5ae92bca01207cdf8ec830785f417437fee4dbb5537a2eeae13c7

Observation 7cb1dfe4-9caa-4660-8e0f-322ee862b86d · outbound

This paper cites Generative Gaussian Splatting for Unbounded 3D City Generation.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Generative Gaussian Splatting for Unbounded 3D City Generation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:46.818968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:46.818968Z digest=sha256:944fe15313ed309598f7da3fa82ae1fff403326affa1c6c1e5bffda759adf390

Observation 8ad4195c-3cd5-4c51-9efb-b631cd4f3f37 · outbound

This paper cites Grid-guided neural radiance fields for large urban scenes.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Grid-guided neural radiance fields for large urban scenes

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:52.184167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:46.895366Z digest=sha256:8489685d78aeed978e4d913bdb21ca4a5265d22b53ec65cfb47885089e7065c7

Observation b760559c-52cb-4306-a908-3e67e5e05eeb · outbound

This paper cites Pointllm: Empowering large language models to understand point clouds.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Pointllm: Empowering large language models to understand point clouds

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:51.883576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:46.938093Z digest=sha256:ca21f7d2623276aa209d13bfa66bb3baa671835155ba7970ee52e081d555ccf6

Observation e3dd9d04-2419-4c03-a6a6-a96cfc73e4c1 · outbound

This paper cites Addressclip: Empowering vision-language models for city-wide image address localization.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Addressclip: Empowering vision-language models for city-wide image address localization

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:51.674509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:47.105618Z digest=sha256:95de8a999e0852bd38c3846feb53058bc5d1832bf5f7e63af8549538c7d98c74

Observation c5b83c66-8f0f-4dff-983b-606d80e0b1ee · outbound

This paper cites Unisim: A neural closed-loop sensor simulator.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Unisim: A neural closed-loop sensor simulator

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:51.420687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:47.217718Z digest=sha256:86fe9811042fe6da3c8c7634c171148d6785a5055e4cab1a6ec84973d6d6c55b

Observation 8a8f6348-3a1f-4b59-8b62-811d51aae34f · outbound

This paper cites Scannet++: A high-fidelity dataset of 3d indoor scenes.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Scannet++: A high-fidelity dataset of 3d indoor scenes

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:51.161037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:47.345334Z digest=sha256:38a7a755292e6fdd96877f13f87ceb1d60e4d68eeef88a821235a022f62fb6a2

Observation 928e6764-3973-4ad1-bf7b-d15f99d7e2e9 · outbound

This paper cites Dogs: Distributed-oriented gaus- sian splatting for large-scale 3d reconstruction via gaussian consensus.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Dogs: Distributed-oriented gaus- sian splatting for large-scale 3d reconstruction via gaussian consensus

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:50.891902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:47.442586Z digest=sha256:598bd14c1f62fb161a384c4bb8ce7cf75bcd05fd95b52bd8e1a6e50d47ef393f

Observation c06fb373-00e8-4647-aca5-bfeb0b7331a2 · outbound

This paper cites PreSight: Enhancing Autonomous Vehicle Perception with City-Scale NeRF Priors.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields PreSight: Enhancing Autonomous Vehicle Perception with City-Scale NeRF Priors

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:49:48.524762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:47.509275Z digest=sha256:d207fe8c83efc0f97c5b232fbbfe7ac1f13ab607930623bb0b9924342e58bb43

Observation 07bb5d59-1330-4260-90f3-f1ce37af3872 · outbound

This paper cites Exploring a fine-grained multiscale method for cross-modal remote sensing image re- trieval.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Exploring a fine-grained multiscale method for cross-modal remote sensing image re- trieval

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:50.727057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:47.551158Z digest=sha256:b681832b44b14e85ece16f4c6fe93610b8b3c84d17f2bf3e1543bd68327b7d4c

Observation c8b06b90-3f35-40f9-bfbd-76a741394e33 · outbound

This paper cites Garfield++: Reinforced gaussian ra- diance fields for large-scale 3d scene reconstruction, 2024.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Garfield++: Reinforced gaussian ra- diance fields for large-scale 3d scene reconstruction, 2024

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:50.489234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:47.607078Z digest=sha256:309166c52894ed79efa6c7c1f79e4a771112fc0e6fa559a02c6964950e44d986

Observation 1444202d-ec7b-4a38-9b77-da8efe02732a · outbound

This paper cites 3DitScene: Editing any scene via language-guided disentan- gled gaussian splatting.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 3DitScene: Editing any scene via language-guided disentan- gled gaussian splatting

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:50.181200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:47.663922Z digest=sha256:f5b95a5a827b2afdd30bfd098de59116c68ce335a5922b0853d6793a6e572e31

Observation 5d435d0e-5c3e-4e41-bda2-42231a60e7fa · outbound

This paper cites Earthgpt: A universal multi-modal large lan- guage model for multi-sensor image comprehension in re- mote sensing domain.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Earthgpt: A universal multi-modal large lan- guage model for multi-sensor image comprehension in re- mote sensing domain

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:49.989182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:47.719996Z digest=sha256:ad930937c3e6ce78f55116f416aceec2f4d9639ac730c58f54474f93f7e51244

Observation cb0f6bdb-ae78-497e-8085-d4022ea2b322 · outbound

This paper cites EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:47.792628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:47.792628Z digest=sha256:66726138383cbf0471fc91627ef9ec63c6bab7cd8fe998acd3b6de46ad73a53a

Observation a6fde006-503c-4e13-a313-7da9fb1a1243 · outbound

This paper cites Efficient Large-scale Scene Representation with a Hybrid of High-resolution Grid and Plane Features.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Efficient Large-scale Scene Representation with a Hybrid of High-resolution Grid and Plane Features

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:49:48.374875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:47.847955Z digest=sha256:2d8d9bb81448d99cd82ac9ad3ea8f4bf2873ee4d88911694ff96960a750acc4d

Observation a7ef6112-0d56-40b4-8e46-add371eadc78 · outbound

This paper cites Aerial lifting: Neural urban semantic and building instance lifting from aerial imagery.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Aerial lifting: Neural urban semantic and building instance lifting from aerial imagery

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:49.770435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:47.900648Z digest=sha256:c0d9b9a459387f5c5c2224046ae98fceb5a174065feeb89c004d390e05013a3a

Observation de21f140-582b-4c75-9827-6bbac2003409 · outbound

This paper cites Rs5m and georsclip: A large-scale vision- language dataset and a large vision-language model for remote sensing.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Rs5m and georsclip: A large-scale vision- language dataset and a large vision-language model for remote sensing

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:49.525950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:47.958787Z digest=sha256:d34d9990e205ebb9fe82dcc5ac0ee223a4fd028e2bd2b32029d62a041127c670

Observation 861b8f96-6aef-48a5-9807-82467594e760 · outbound

This paper cites Mutual Attention Inception Network for Remote Sensing Visual Question Answering.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Mutual Attention Inception Network for Remote Sensing Visual Question Answering

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:49.261423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:48.037601Z digest=sha256:2cecda84fd41703d090080f0cf95bd2095570fec200ac1b1b2166604f01c6d29

Observation e9afdf7b-ffbb-44f9-8cac-136746af8ab1 · outbound

This paper cites Drivinggaussian: Composite gaussian splatting for surrounding dynamic au- tonomous driving scenes.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Drivinggaussian: Composite gaussian splatting for surrounding dynamic au- tonomous driving scenes

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:49.088711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:48.120675Z digest=sha256:7da117829ec9c9d173d86533a585e5a974c1690a8635d53944d4f71fb791e75a

Observation 067fb573-df17-4484-9755-f991eb3f2950 · outbound

This paper cites Towards vision- language geo-foundation models: A survey.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields Towards vision- language geo-foundation models: A survey

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T21:49:48.162222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:49:48.162222Z digest=sha256:3a5cf3d6db9ca2db573b6571beacd3e8a7d29278f68a47b6e62676b9ce5aaf56

Observation 665b8c4d-6538-4db8-9d99-a80c8c9f3ed5 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:48.941651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:48.227457Z digest=sha256:202466757ac144a331e7689afd6f811de1958bc963cfaa55642672d4b777d19d

Observation 0eddabe3-4fb5-4160-a229-5338d77d636e · outbound

This paper cites 'yes' if {ANSWER1} < {ANSWER2} else 'no'.

GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields 'yes' if {ANSWER1} < {ANSWER2} else 'no'

Reference 2023

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T21:49:48.805064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:49:48.230001Z digest=sha256:d6b0f2a15a8c28207c7f3ae683aff28a7b8d3203102b90b4f05f719e2f752f1f

Pith citing papers

No inbound Pith citation observations are available.