Pith. sign in

Paper Citation Record · LEDGER

Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2309.00615.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.00615 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T12:50:01.034749Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:09:29.321784Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c7239216-2861-4a65-ace9-fbe23b8d32f4 · inbound

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention cites this paper.

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 109

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T23:07:42.622601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-14T23:07:42.245641Z digest=sha256:989697844a66db832278b165dbd8f2034cbaa04b621541165fd15b85b66b3f28

Observation 2b2ea280-db75-4de4-bd15-237819b722e0 · inbound

SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models cites this paper.

SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:03:26.883172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T03:03:26.723464Z digest=sha256:9b6c72083bf7f9d9dcfeb6ad250968fc46891c3d7180bd8fb85be6ca411662ec

Observation 63e58c97-9553-4247-933f-e64629c117d9 · inbound

MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? cites this paper.

MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T01:29:30.121405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T01:29:30.032408Z digest=sha256:4a71d1bd88cae971d40019869cec1c49a3ca4ed296a396b5752d98209ed20c31

Observation 080577c8-4f50-4d7f-a6ab-4f3b2ac2b91e · inbound

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models cites this paper.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:54.075378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:35549611af1e5659831605b1bd31ca6ff199d8e74782690ad8af3c5cffff1fa5

Observation ee4d18be-d8a8-4d3a-8433-57f5ae56d34d · inbound

LLaVA-OneVision: Easy Visual Task Transfer cites this paper.

LLaVA-OneVision: Easy Visual Task Transfer Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:23:49.552473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:abecc3b0b3ab23bca4dd288e128cc574f8cfa879eabcb18043820f58fd921444

Observation 6b4cb0ed-a327-4aff-827a-e27a4f09a8dc · inbound

Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark cites this paper.

Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:05:47.810561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-23T20:03:38.336841Z digest=sha256:5b44aadf1365bd16565e39bec486b8670bfad72067f3c1af4afda9211f7ee53f

Observation fb9951c6-01da-41c6-b870-43e634eac887 · inbound

SVL: Spike-based Vision-language Pretraining for Efficient 3D Open-world Understanding cites this paper.

SVL: Spike-based Vision-language Pretraining for Efficient 3D Open-world Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:14:31.402477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-22T02:12:29.058875Z digest=sha256:35b3311b4914e1997ff871b50853efb6dc828f5438bb5630d06dc7fbcbd4d393

Observation 84f5614f-93ca-4b8c-8136-58b77355c6a4 · inbound

Calibrated Multimodal Representation Learning with Missing Modalities cites this paper.

Calibrated Multimodal Representation Learning with Missing Modalities Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:10:22.660828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T22:08:07.217659Z digest=sha256:69d98f7e784153c24cb616b8552587487c5c70d80698b070dc4673bf46fbc9ac

Observation e61121aa-809c-44e4-a755-e5b43e9b7d83 · inbound

Rethinking Multimodal Few-Shot 3D Point Cloud Segmentation: From Fused Refinement to Decoupled Arbitration cites this paper.

Rethinking Multimodal Few-Shot 3D Point Cloud Segmentation: From Fused Refinement to Decoupled Arbitration Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T12:50:01.034749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:50:01.034749Z digest=sha256:16c42a4fc13303b451f8191b301f90921df68603db095ca77448bf1eaaf267fa

Observation 26024ed0-c9be-4d31-be8a-7c3e922ab0e1 · inbound

Pointy - A Lightweight Transformer for Point Cloud Foundation Models cites this paper.

Pointy - A Lightweight Transformer for Point Cloud Foundation Models Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T13:30:01.609235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T13:28:23.934522Z digest=sha256:efa0cd867fdd7008ae9bbb64be7c973b0cb8fef6fdc985d31d51667d7db2a8fd

Observation 860ced58-9fcd-4dd9-a5c6-08df71a734a3 · inbound

Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding cites this paper.

Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:30:22.462415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T09:30:03.668178Z digest=sha256:07255c9ba1b90f7b78c41c0fa3f9489a2daac5c47416b19e5c5450cb5be86200

Observation fc0d30ca-9ace-47ee-aae3-8a25e85ce2a6 · inbound

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding cites this paper.

Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T17:54:04.904102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:54:04.904102Z digest=sha256:80fbd52f81e1c4a47f324f62a541b6eb2836ce75dbc04cfbf39663feaf87ebed

Observation 2a95df81-557a-4628-91a3-784c68deacc3 · inbound

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM cites this paper.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.409437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:46d8e9aad8083518d27d501c112bfc0771e1942d33587b0f045ef92f3555d03f

Observation 5ec2897e-d7c3-4518-a28f-f3a1f3b05512 · inbound

RGB-Pointmap Pretraining for Unified 3D Scene Understanding cites this paper.

RGB-Pointmap Pretraining for Unified 3D Scene Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T21:23:17.900121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T21:19:49.421653Z digest=sha256:1cf30dfad4368e6bd5c74c4b67638d8e53d2cdf60d15121128bbc9e01100a922

Observation 8ad60e09-435a-4793-aa0d-2884ed868abf · inbound

Efficient3D: A Unified Framework for Adaptive and Debiased Token Reduction in 3D MLLMs cites this paper.

Efficient3D: A Unified Framework for Adaptive and Debiased Token Reduction in 3D MLLMs Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:48:11.404582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T19:45:33.950587Z digest=sha256:d597260f8688758dbc963a2df09d7b08a09bafb81746ba8d0ca20e8592ac64b8

Observation bb3797ac-0573-40bb-893f-d23c12f67e7d · inbound

Beyond Surface Artifacts: Capturing Shared Latent Forgery Knowledge Across Modalities cites this paper.

Beyond Surface Artifacts: Capturing Shared Latent Forgery Knowledge Across Modalities Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:30:59.344899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:06:06.896112Z digest=sha256:29ac3a9cb990f037b46d9870d66ecff27703ddce348e68cbea57bbf2898ca683

Observation 22b41bc1-497c-41a9-90d5-cfed256071a6 · inbound

Reinforcing 3D Understanding in Point-VLMs via Geometric Reward Credit Assignment cites this paper.

Reinforcing 3D Understanding in Point-VLMs via Geometric Reward Credit Assignment Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:04:18.054828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-09T23:00:14.031709Z digest=sha256:f32756b08d79b0912b8e0c96081b2faec8093e4492c16a8ce42747ebea70acb4

Observation 41006236-384f-48d0-967e-987897c84767 · inbound

SGSoft: Learning Fused Semantic-Geometric Features for 3D Shape Correspondence via Template-Guided Soft Signals cites this paper.

SGSoft: Learning Fused Semantic-Geometric Features for 3D Shape Correspondence via Template-Guided Soft Signals Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:53:15.141947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T11:48:55.446684Z digest=sha256:90938000203a700b0d830092cf9465869f2dca1c657fb287626adb9824bc1df3

Observation bb6a6b9a-3818-4e0a-a863-38a7bfce053d · inbound

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning cites this paper.

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:18:57.763666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-27T01:23:40.564561Z digest=sha256:fda8f31ff254d6fcdfdcd2f4a91ec0caee1464d9f5dc633f6dbf09c15c413b50

Observation e7e808af-abb6-456d-a443-0ebefd3e80f8 · inbound

Occ-VLM: Occupancy Grounded Vision Language Model for Indoor Scene Understanding cites this paper.

Occ-VLM: Occupancy Grounded Vision Language Model for Indoor Scene Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T02:59:25.799065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-26T18:40:20.588652Z digest=sha256:2907843dc89c9a512fd6eea570eef250d1801bebdfb1eab3ca09e2c874b82291

Observation 083afcaf-f843-4667-976e-a20b19ff9d03 · inbound

3D-PLOT-LLM: Part-Level Object Tokens for 3D Large Language Models cites this paper.

3D-PLOT-LLM: Part-Level Object Tokens for 3D Large Language Models Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:09:29.323841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-26T18:25:55.011125Z digest=sha256:9602b60ca8f845d86ec8bb46e468cc0a6517dc13f6aa7cae7d03d62ed67e24c8

Observation cc7bc106-3e0b-4cde-94ca-d2f97f99e77a · inbound

MV-GEL: Language-Driven Multi-View Geometric Entity Localization on Meshes cites this paper.

MV-GEL: Language-Driven Multi-View Geometric Entity Localization on Meshes Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:05:40.451224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-01T05:56:47.849311Z digest=sha256:099d0dd252e2a40599244e880e64dea1fff3a81230aa0e2231089989bce9cb90

Observation 85be3b90-7c38-49e4-8bab-740661559d48 · inbound

MV-GEL: Language-Driven Multi-View Geometric Entity Localization on Meshes cites this paper.

MV-GEL: Language-Driven Multi-View Geometric Entity Localization on Meshes Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T09:23:10.452175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:23:10.452175Z digest=sha256:47f783a3638e0c5febedde1187663673be02fd6abb254536632c011c4b6d4f07

Observation d1acf887-f64c-4ed8-a250-ec3305992d20 · inbound

SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video cites this paper.

SpaceEra++: A Unified Framework Towards 3D Spatial Reasoning in Video Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:18:37.260308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T16:16:41.412451Z digest=sha256:5cde5a93bae7abaf9cd5bacff7291f4c16aaed31f1030e06123f10cebad09e22

Observation 131fa64a-d148-48f8-88e0-f269398cb505 · inbound

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation cites this paper.

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T06:23:28.101393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:23:28.101393Z digest=sha256:f87e3e0ae3db2a37630d7a971496205f0e6c86c83a6588ddad4e89573223d26d

Observation 4ba8b9c3-c4ab-43e9-9f7c-f875897ca14e · inbound

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding cites this paper.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.446109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.446109Z digest=sha256:85d18542a98fcafbfc6360a843778af05b6751ad6743a08658d4b299a264e1ca