Pith. sign in

Paper Citation Record · LEDGER

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models

As of 24 July 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2605.25901.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.25901 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T22:47:05.747852Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-24T06:31:00.690269+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 621ec68a-bcc2-4f28-9af0-710e32d3bff8 · outbound

This paper cites Gˆ 3-lq: Marrying hyperbolic alignment with explicit semantic-geometric modeling for 3d visual grounding,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Gˆ 3-lq: Marrying hyperbolic alignment with explicit semantic-geometric modeling for 3d visual grounding,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:22e2783071add3e80fc891b7dc82eaf5413d64c536ec3bd33e4fb3f8fb9027e2

Observation c9da0c0c-1e08-4048-a7c0-9b72e9b54f3f · outbound

This paper cites Multi-branch collaborative learning network for 3d visual grounding,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Multi-branch collaborative learning network for 3d visual grounding,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:a3baa79df7583125783a8e117a56b19145549889004dc408d49e89d045ce4688

Observation d9ab67e8-a0f2-438c-95ec-1d65b3489d65 · outbound

This paper cites Chat-scene: Bridging 3d scene and large language models with object identifiers,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Chat-scene: Bridging 3d scene and large language models with object identifiers,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:38d1158edd75e83ae61160a06b1854f4e44c2337142f786b006a5884e0b7e699

Observation 51e219cd-d1ab-4265-add6-d9737eae432b · outbound

This paper cites Video-3d llm: Learning position-aware video representation for 3d scene understanding,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Video-3d llm: Learning position-aware video representation for 3d scene understanding,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:c03e9bc48651d2390334084c0b986185ff0ae6f821aedf47c91fc7782bf4de93

Observation 7102f1d4-a20b-4d15-b2d2-ae916c3d3072 · outbound

This paper cites Llm-grounder: Open-vocabulary 3d visual grounding with large language model as an agent,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Llm-grounder: Open-vocabulary 3d visual grounding with large language model as an agent,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:41ef8ceea49088a63557719f95dcab6149f01b86a6dece84b42804eb0cd7344f

Observation eb61c667-2b73-4a02-8300-b16713cbaef3 · outbound

This paper cites Visual programming for zero-shot open- vocabulary 3d visual grounding,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Visual programming for zero-shot open- vocabulary 3d visual grounding,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:1f4e4c5ee11b98e051dc8feac03c03cdbde52766c7fc6fb1f492c950586d8a3d

Observation 8c1e8d07-ce2a-4b31-83dd-007933d773e5 · outbound

This paper cites Vlm-grounder: A vlm agent for zero-shot 3d visual grounding,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Vlm-grounder: A vlm agent for zero-shot 3d visual grounding,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:aa99586dae746abf6417ac514f1a8be1b2608a7277a9baa54eafd0be961b1660

Observation a94e15b2-98c4-4047-a865-3cd6849ea469 · outbound

This paper cites See- ground: See and ground for zero-shot open-vocabulary 3d visual grounding,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models See- ground: See and ground for zero-shot open-vocabulary 3d visual grounding,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:a1bc6fded2077c3895b5e57127efcf4583d0baf609aaa60d4fbc91b74d23341e

Observation 4bd7f0ab-0d5b-43f0-8e52-8c1aa303fa51 · outbound

This paper cites Solving Zero-Shot 3D Visual Grounding as Constraint Satisfaction Problems.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Solving Zero-Shot 3D Visual Grounding as Constraint Satisfaction Problems

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:54:01.240702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:38cbe9ac7b1e616ed1925f40c74bed470bae1f5cd6f4772cd8d19d646c8f736c

Observation bd0b770d-1351-4f08-998d-733a7bf42cf4 · outbound

This paper cites Sort3d: Spatial object-centric rea- soning toolbox for zero-shot 3d grounding using large language models,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Sort3d: Spatial object-centric rea- soning toolbox for zero-shot 3d grounding using large language models,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:e69c52ad93dd6737bd868fe2af6bf84ccd1ecbf162efaa736dc534fcf26cb9db

Observation 5abcaaca-cd13-4feb-9078-2603b60edad4 · outbound

This paper cites Transcrib3d: 3d referring expression resolution through large language models,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Transcrib3d: 3d referring expression resolution through large language models,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:a8e0d7b363b5c759d45d49432475f4a7d8216fd4b01d7891936e94b9a6746e00

Observation 1a07273b-3ba3-4af3-ad29-674b8084aa9e · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural lan- guage,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Scanrefer: 3d object localization in rgb-d scans using natural lan- guage,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:f4100c5c81d7368479edf53d2ddf8a42f31f998eeb945d4017525eb60f590b13

Observation ba5ea7fa-6a64-4c12-84f1-d9db03d5b78c · outbound

This paper cites Referit3d: Neural listeners for fine- grained 3d object identification in real-world scenes,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Referit3d: Neural listeners for fine- grained 3d object identification in real-world scenes,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:61e53aef335967011050d4b8dfc9ee4d746f51f0f3e187ca8819a9ae6d7041b5

Observation 38b20b93-1502-4d02-b450-39428e94ab7e · outbound

This paper cites Four ways to improve verbo-visual fusion for dense 3d visual grounding,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Four ways to improve verbo-visual fusion for dense 3d visual grounding,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:a481509d5d198b88d3548fbbcc9e7aa5ce9e798b878af382e81f1529f07dd54b

Observation bb040efe-7a4b-43af-80fa-113c0f2b131c · outbound

This paper cites OpenMask3D: Open-Vocabulary 3D Instance Segmentation.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models OpenMask3D: Open-Vocabulary 3D Instance Segmentation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:54:01.231727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:2fc6f9fcb79cc813c2626764dea18a5f429bb40b61f42e543d591f47095da3ee

Observation 0ff0fb14-1a73-44f0-b6cc-89ca21138747 · outbound

This paper cites Isbnet: A 3d point cloud instance segmentation network with instance-aware sampling and box-aware dynamic con- volution,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Isbnet: A 3d point cloud instance segmentation network with instance-aware sampling and box-aware dynamic con- volution,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:e4501e68b3e9916a82aa7b241de1a7d1c1b9de6c442bd38e0a403cd426be6b32

Observation 45358fbe-78bc-4736-8fc6-78d01825bd88 · outbound

This paper cites SAM3D: Segment Anything in 3D Scenes.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models SAM3D: Segment Anything in 3D Scenes

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:54:01.237749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:7e516ab6b2fbf38685e9c02f3f98c230ca49f4990b585c31348b5a1bff85bfe6

Observation cf86a4f2-e806-457b-9169-53279f855ad0 · outbound

This paper cites Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Open3dis: Open-vocabulary 3d instance segmentation with 2d mask guidance,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:45bf49d5664c6757d888ae18bd7c7f1863cb86805fd45337390ea9b294d34781

Observation 49e580fe-5d5b-4830-9e83-65bb7bb6026e · outbound

This paper cites Any3dis: Class-agnostic 3d instance segmentation by 2d mask tracking,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Any3dis: Class-agnostic 3d instance segmentation by 2d mask tracking,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:11a4b717c5f6699ec289a2ace033038aac79381da34205523277128565a877d6

Observation 4e1d08be-0b11-4443-b714-081085a2e5dc · outbound

This paper cites Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:99b249fcf6e94b6b241c806b8f0ce740eeacd40d6f39cc78de1d04587cb79015

Observation e75f00bd-bd61-424d-a91c-acb2c5eb3f5f · outbound

This paper cites ConceptFusion: Open-set Multimodal 3D Mapping.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models ConceptFusion: Open-set Multimodal 3D Mapping

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:54:01.246558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:650a78c34299ef37ad0004053fa32b2e036d3d0fec5d2beb78a90f543a9aa7bb

Observation 331a537f-d8b2-4772-becc-1fca06193087 · outbound

This paper cites Phygrasp: Generalizing robotic grasp- ing with physics-informed large multimodal models,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Phygrasp: Generalizing robotic grasp- ing with physics-informed large multimodal models,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:5c68e43dc47741a4bda0da7c9f12fa0fe38acd05fd267cc17d0c7f986302e9ea

Observation af2ec1cf-a924-498c-8813-9c69a17fe1e3 · outbound

This paper cites Zero-shot object navigation with vision-language models reasoning,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Zero-shot object navigation with vision-language models reasoning,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:027caad85bc86b328061ec2465ed9bf75cbaf9893ad2f32e3b992176e60615e4

Observation 3ec2d34f-b841-446d-87f2-776521f3067b · outbound

This paper cites Iref-vla: A benchmark for interactive referen- tial grounding with imperfect language in 3d scenes,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Iref-vla: A benchmark for interactive referen- tial grounding with imperfect language in 3d scenes,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:12f2365d3df4c7d0449ccc5912576bc0b5e67860dbfa47df98e4e6b5104a3f4e

Observation 561bfa1f-4a12-4808-94b8-b6571e253da6 · outbound

This paper cites Sceneverse: Scaling 3d vision-language learning for grounded scene understanding,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Sceneverse: Scaling 3d vision-language learning for grounded scene understanding,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:82818bb64d36899da17a9658b67ddf86f8935e2a482c923c3f5240367d751a37

Observation 5bf9c083-32d1-4bf2-bb20-047d1d535f0c · outbound

This paper cites Mask3d: Mask transformer for 3d semantic instance segmentation,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Mask3d: Mask transformer for 3d semantic instance segmentation,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:6eebc6474463af1e2c11129f320a46dfdf3297ebe5de99c135385009f56429a9

Observation f01fbfbe-c7a8-4faf-8764-197323671bb7 · outbound

This paper cites Qwen3 Technical Report.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Qwen3 Technical Report

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:54:01.243392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:dc21fa9f315479bf207f3807deadd8fd26522977e97a89897efaf5abf814109c

Observation 0c0e2192-34ab-4268-bdf8-e0225399c842 · outbound

This paper cites Using ollama,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Using ollama,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:56e07b1e372efece1470c89e0a115d2a7e929d03bbdbe101280f88a4797265a7

Observation aa9cb1c7-2d56-470b-80c0-b7049d59ccc1 · outbound

This paper cites Chase,Langchain,https : / / github.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Chase,Langchain,https : / / github

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:7e7cfe851b53202570b24cb45473a9e234f9c1b2d9aeec6556d62d6134882960

Observation 10454cfd-1cdb-4d81-94e9-768de310711f · outbound

This paper cites GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:54:01.235024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-24T06:31:00.690269+00:00.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:337a95f154d06e4abb081dd621e5879153a10da07a9b4f4438c476e14efe327f

Observation 27652e4b-f858-40ac-886e-d13de3c027a6 · outbound

This paper cites Mikasa: Multi-key-anchor & scene-aware trans- former for 3d visual grounding,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Mikasa: Multi-key-anchor & scene-aware trans- former for 3d visual grounding,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:725b00e1d746a856646e1ae165a3b2f520446fb93d4ee753eb55111a41c7e011

Observation ab0266f1-8843-4483-994b-352c5dcba901 · outbound

This paper cites Language conditioned spatial relation rea- soning for 3d object grounding,.

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models Language conditioned spatial relation rea- soning for 3d object grounding,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-29T22:47:05.747852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T22:47:05.747852Z digest=sha256:648ecd556ba8d0fe25ce1fce14fdaec089435dad48c8f28d27e2f51429bcf4fe

Pith citing papers

No inbound Pith citation observations are available.