Pith. sign in

Paper Citation Record · LEDGER

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments

As of 15 August 2026, this Paper Citation Record lists 10 of 10 outbound references and 0 inbound Pith citation observations for arXiv:2602.18527.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.18527 v3

Coverage vector

measured 10 of 10 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T22:08:40.534320Z

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

10 of 10 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a94c27af-931a-4fff-bf49-86ccdf60e939 · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments Grounded 3D-LLM with Referent Tokens

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T22:08:39.759563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:08:39.759563Z digest=sha256:5cc2a3cfe395b64d7476323222fc6b408f7e085328030aec6245d3f73249546a

Observation 46e24356-6c6d-4f8f-a228-c091bdbfdbd2 · outbound

This paper cites video-SALMONN 2: Captioning- enhanced audio-visual large language models.arXiv preprint arXiv:2506.15220,.

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments video-SALMONN 2: Captioning- enhanced audio-visual large language models.arXiv preprint arXiv:2506.15220,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T22:08:40.139307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:08:40.139307Z digest=sha256:c2bad2f7de1b19c3cc7783a5375b4d756fa1cedfaaf6cd4f6a015fd1386f16eb

Observation 1a2f57f0-45d7-4771-bc89-eca22ec1c847 · outbound

This paper cites N3D-VLM: Native 3D ground- ing enables accurate spatial reasoning in vision-language models.arXiv preprint arXiv:2512.16561,.

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments N3D-VLM: Native 3D ground- ing enables accurate spatial reasoning in vision-language models.arXiv preprint arXiv:2512.16561,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T22:08:40.268567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:08:40.268567Z digest=sha256:3e26c94eebf594377d60acff52ffa8849d536c06e690411e8670a04081fbfc06

Observation c40a3b6a-1571-430d-8a5a-b23d10798b52 · outbound

This paper cites Qwen2.5-Omni Technical Report.

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments Qwen2.5-Omni Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T22:08:40.395661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:08:40.395661Z digest=sha256:588900aff07ec36907a684c325df1452a95f8f78c649633c72da4dffa13380d3

Observation 7736e1c7-74bc-459d-84f9-2d0ce48969e2 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T22:08:40.534320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:08:40.534320Z digest=sha256:a16613e28e5e999654481366b1a741479c3e74224024bf22371282b319f4e9e5

Observation 48b73ff5-a861-434c-9f81-a910cb1ca327 · outbound

This paper cites GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models.

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-02T22:08:40.003799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:08:40.003799Z digest=sha256:72966878445a471f35dbb141cb645e05b627498a1d38878e74b01fc82f294542

Observation ff301b4e-f09b-4ef5-a34b-8ce43bfdbda7 · outbound

This paper cites SpatialVLM: Endowing vision- language models with spatial reasoning capabilities.

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments SpatialVLM: Endowing vision- language models with spatial reasoning capabilities

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T22:08:39.581432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:08:39.581432Z digest=sha256:f6271e2f99f1feb4287a6d31dffb9d497aedfd000b56d6651111aefee60be333

Observation 66a0e07c-dd83-4b94-9929-b405c103bbd6 · outbound

This paper cites Qwen3-VL Technical Report.

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments Qwen3-VL Technical Report

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T22:08:39.292349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:08:39.292349Z digest=sha256:a096c35b724317794061dbaf144c751abf5d35864ed4d3f3c939143f4e94aabd

Observation 70193c6f-cc14-436f-8edf-b32e7c284204 · outbound

This paper cites Seed1.5-VL Technical Report.

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments Seed1.5-VL Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T22:08:39.921774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:08:39.921774Z digest=sha256:99709afe3ebb2156beedf9808262c1fdb72191d54b1c0d9f66e95cb2f7f33fbc

Observation dd0715a1-8cf8-433d-adc8-c8cf6fa447ab · outbound

This paper cites an unresolved cited work.

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T22:08:39.399844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:08:39.399844Z digest=sha256:762bfbbc79d17dc00030ac463014043b5166fe186706ccc2d3bb2deb98165ff1

Pith citing papers

No inbound Pith citation observations are available.