Pith. sign in

Paper Citation Record · LEDGER

PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding

As of 6 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2604.15770.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.15770 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T08:31:08.720736Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 95d7f15e-7b34-4fc8-89ce-fd688f080e09 · outbound

This paper cites Openscene: 3d scene understanding with open vocab- ularies.

PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding Openscene: 3d scene understanding with open vocab- ularies

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:24:17.883595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T08:31:08.720736Z digest=sha256:9882d9aef7eaaa85c24ca2e592a5f31e0c2cf855e64766cb32668fbc9c1d33e9

Observation 81a32c7f-77cd-4455-8e30-fce0f5c29963 · outbound

This paper cites Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning.

PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.796763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T08:31:08.720736Z digest=sha256:2d05139d2c199feab57d4765d0e38d9bdf14edac52151c34055066081bb1d25e

Observation f4645d52-f4cf-4d84-97ee-634790cc8720 · outbound

This paper cites Learning transferable visual models from natural language supervision.

PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding Learning transferable visual models from natural language supervision

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.801196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T08:31:08.720736Z digest=sha256:aa058d91dd6cff38752ce848edb72862e7a5a0df2b41e13b590544d4e6cfbaee

Observation 21df3d4c-ad52-49bb-bc25-88f48338b263 · outbound

This paper cites Conceptfu- sion: Open-set multimodal 3d mapping.

PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding Conceptfu- sion: Open-set multimodal 3d mapping

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.804996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T08:31:08.720736Z digest=sha256:9a67900a9aa151f0ca856c2cc33e103321dfdd44a8a09971a65a6af25148c2aa

Observation d03d22b9-4650-4bc7-a8f7-60ee5d9673dd · outbound

This paper cites Openmask3d: Open-vocabulary 3d instance segmen- tation.

PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding Openmask3d: Open-vocabulary 3d instance segmen- tation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.792397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T08:31:08.720736Z digest=sha256:bfb708977e85e9b66ab060b8f8dfb49aea18aef2365d52465b116b10171e42a1

Observation 26af8c7b-e067-49ae-8637-9fb5005b78bf · outbound

This paper cites Am-radio: Agglomerative vision foundation model reduce all domains into one.

PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding Am-radio: Agglomerative vision foundation model reduce all domains into one

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.812774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T08:31:08.720736Z digest=sha256:b078f6f71914b6199e49075e84428dcdae45a693c11ee6d17012e5042264a86e

Observation 774ddae1-36b6-4640-afdf-05aec72ed9ee · outbound

This paper cites RayFronts: Open-Set Semantic Ray Frontiers for Online Scene Understanding and Exploration.

PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding RayFronts: Open-Set Semantic Ray Frontiers for Online Scene Understanding and Exploration

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:53:05.074208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T08:31:08.720736Z digest=sha256:859b48baa00dfd90f514e4488c701ee7d86d515bf6ad7a9ec43d9f4afa9a9897

Observation 3d48a868-d852-412a-9a37-8899c2785049 · outbound

This paper cites Inst3d-lmm: Instance- aware 3d scene understanding with multi-modal instruction tuning.

PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding Inst3d-lmm: Instance- aware 3d scene understanding with multi-modal instruction tuning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:14:18.808815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T08:31:08.720736Z digest=sha256:ba253d86df2fde228754788c83df33e070ade05fb039236669d0b64fa8bbeab5

Observation bc3bb080-353b-4ff5-838f-ca13d9b65be0 · outbound

This paper cites Dinov2: Learning robust visual features without supervision.

PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding Dinov2: Learning robust visual features without supervision

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:24:17.903751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T08:31:08.720736Z digest=sha256:12db9cb21c189133fbf912d5387cf91810f87bc92ffea0817bcaad59509ff0d7

Observation dc7f8c59-c120-4bf2-9a4f-29bf249840aa · outbound

This paper cites Radiov2. 5: Improved baselines for agglom- erative vision foundation models.

PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding Radiov2. 5: Improved baselines for agglom- erative vision foundation models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:24:17.889516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T08:31:08.720736Z digest=sha256:02f14950c0444cff7f6dcdc55d008e5b7e689863861cdeb93460f37e0c14b42f

Observation 48c5a499-f737-4f23-a7b6-66dc01d1d24f · outbound

This paper cites Segment anything.

PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding Segment anything

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:24:17.894879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T08:31:08.720736Z digest=sha256:7ae71c5b8013edcfb3cf0b00f9d1949599d2d4da41c0ade6393952fccd625df2

Observation efdeb7dc-7c99-44a5-93f3-b9417006bf7d · outbound

This paper cites Clip-fo3d: Learning free open-world 3d scene representations from 2d dense clip.

PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding Clip-fo3d: Learning free open-world 3d scene representations from 2d dense clip

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:24:17.899378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T08:31:08.720736Z digest=sha256:960a258ef9aba40c1fa521e8a6ffeeb241e63a48eb5172817d6443f92dd5f1a5

Observation 80990d33-5ac8-49a6-8544-7b5a516084a8 · outbound

This paper cites Neural compression-based feature learning for video restoration.

PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding Neural compression-based feature learning for video restoration

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:24:17.924278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T08:31:08.720736Z digest=sha256:af9020b9c8c31545359985b26a22dbce08918414def654399083033b7aeaca04

Observation 1ac679b6-64d8-48f0-87fb-fc50737001a6 · outbound

This paper cites Fully sparse 3d occupancy prediction.

PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding Fully sparse 3d occupancy prediction

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:24:17.961271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T08:31:08.720736Z digest=sha256:66a82c63545d6523351740c671e0c0a34ef9525599a2c7c60f6eb8ebb1657847

Observation 33f312bb-e2a2-42e0-861b-88d84e9ad648 · outbound

This paper cites Kimera: an open- source library for real-time metric-semantic localization and mapping.

PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding Kimera: an open- source library for real-time metric-semantic localization and mapping

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:24:17.934916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T08:31:08.720736Z digest=sha256:478432a5d06c7b86cc128e785b5dd6bee61988d795ebb36377a02d8e1f96a8f5

Observation 6fce7cad-7088-4f2f-a82b-ab6c4279fff6 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:24:17.930626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T08:31:08.720736Z digest=sha256:ed62dcf5599438a73a383bf43a3455976d5f1d2a122c6a2803897cc94ff8e13d

Observation 385430ba-00a5-4d53-97be-84058c86d721 · outbound

This paper cites Scene parsing through ade20k dataset.

PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding Scene parsing through ade20k dataset

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T00:24:17.968513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T08:31:08.720736Z digest=sha256:2d69a024cc571bd119b57ce07f050188fe6d1607547d1a0f6eae072c697b8273

Pith citing papers

No inbound Pith citation observations are available.