Pith. sign in

Paper Citation Record · LEDGER

DenseMLLM: Standard Multimodal LLMs for Dense Prediction

As of 6 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2602.14134.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.14134 v2

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T23:23:35.957580Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-08T01:54:30.649092Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T02:04:26.424891Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 96f62d90-8e2e-41b6-ad2d-0f6fed872828 · outbound

This paper cites RLE string.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction RLE string

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:35.463345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:35.463345Z digest=sha256:05c931ca0a2099822351d8c15d384ad77dccc1b5fb3b74f0c687e79c1df06e51

Observation 658d868e-8250-4af4-8bc0-9ce006f3c349 · outbound

This paper cites an unresolved cited work.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:35.718786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:35.718786Z digest=sha256:8c1087cab82db05bff47858a745ddc1c9988348fe2bb0e555f0e4f193a4805db

Observation f46f50ed-9152-4b56-91ae-a312b9e0d20d · outbound

This paper cites S., and Lin, M.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction S., and Lin, M

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:34.302345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:34.302345Z digest=sha256:11e14084f42fb452e4fdf04fe2b36bcc5fa0e902e76a55ae8f0e59e6529b226e

Observation a15d26ea-df99-4243-8734-283e6f38b8eb · outbound

This paper cites Swinmtl: A shared architecture for simultaneous depth estimation and se- mantic segmentation from monocular camera images.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction Swinmtl: A shared architecture for simultaneous depth estimation and se- mantic segmentation from monocular camera images

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:34.473897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:34.473897Z digest=sha256:760ba17a595381acf5ec0bafa739e47e50e54d5ec9a36803b3aa1159597fee3f

Observation 75207e2f-7548-48ec-9a5b-1503e99d3947 · outbound

This paper cites Ufo: A unified approach to fine-grained visual perception via open-ended language interface.arXiv preprint arXiv:2503.01342, 2025a.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction Ufo: A unified approach to fine-grained visual perception via open-ended language interface.arXiv preprint arXiv:2503.01342, 2025a

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:34.644056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:34.644056Z digest=sha256:61375d967486c09a64170e3f5f3a2a4639e02d17f3184f241bacdfdddf119684

Observation 52b6dc11-bfab-4cda-ac4c-669226af52ec · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:34.805229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:34.805229Z digest=sha256:7f98b5c5bf4675664d98141a7cdc0de3962869b8ff41c5f4282f3de7eb13b8aa

Observation bd6dfdfe-4afe-437f-a1a5-f4490fc688d6 · outbound

This paper cites VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:34.980496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:34.980496Z digest=sha256:0f92925163cfb454812ed05b91f9917f6aac4878fc01357ba1b2f97e7cc05b3c

Observation 6645d106-7b66-4803-9bb6-a7c6d0e446d4 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:35.129957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:35.129957Z digest=sha256:ecc9e226f14d07e2aec1ca58af02de077fe6aaaab047b7bcdd366c53477157b3

Observation e52d6eb0-fe4a-4304-ad84-425e73b3fbdd · outbound

This paper cites Qwen2.5 Technical Report.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction Qwen2.5 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:35.241079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:35.241079Z digest=sha256:3072dcdd32e8d6a0691aef0265dcc6d40a1513a0f8fdacb960a03895a4a6a052

Observation 94644f63-9640-493d-983b-78c3f719e063 · outbound

This paper cites Visual representation alignment for multimodal large language models.arXiv preprint arXiv:2509.07979,.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction Visual representation alignment for multimodal large language models.arXiv preprint arXiv:2509.07979,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:35.288627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:35.288627Z digest=sha256:6728072a0c8a2b812bd4f5e0ad1e5f7e2e297468a8ef5df792f5e8be98145153

Observation af58203f-071f-47fd-b286-8d4ed74fdd03 · outbound

This paper cites Semantic Segmentation for the Open World.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction Semantic Segmentation for the Open World

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:35.634650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:35.634650Z digest=sha256:a13bdd42889b9f728a02c55f684558bc1784271e9c00f7c0f2fbdc1f69a69507

Observation f7980be9-7a39-448f-b585-91aa30436ebd · outbound

This paper cites RLE string.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction RLE string

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:35.775714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:35.775714Z digest=sha256:4dc79e29c36279a3e43d0ffa407f55a22621badde2bffad4f34d6063ee8603e6

Observation 721c078c-2fcd-468f-83c6-d0306fc8ebe2 · outbound

This paper cites During testing, the predicted values need to be de-quantized to obtain the real depth; otherwise, only relative depth is obtained.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction During testing, the predicted values need to be de-quantized to obtain the real depth; otherwise, only relative depth is obtained

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:35.957580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:35.957580Z digest=sha256:a48af22e000ab683dedc8ede79a6cd787c4022c926cff2a7b170ac8c5f12ec9b

Observation 1e295596-64b2-4160-a1f5-c77d0116f190 · outbound

This paper cites During testing, we dequantize to the actual depths and exclude invalid depths.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction During testing, we dequantize to the actual depths and exclude invalid depths

Reference 1000

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:35.913684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:35.913684Z digest=sha256:32464d839f08130daabcb94351e376f8b973c119d045159791f19c60979de641

Observation e449bbf3-80e5-470f-b876-fd81be049334 · outbound

This paper cites Token activation map to visually explain multimodal llms.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction Token activation map to visually explain multimodal llms

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:33.922033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:33.922033Z digest=sha256:8a2427afa5a825078c8101273abe0c0958b76362c6188f4b14f2632a5b4ba488

Observation 1451865a-7c9a-4791-9a06-15a32fff1dfe · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction DINOv2: Learning Robust Visual Features without Supervision

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:34.118285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:34.118285Z digest=sha256:fc9a2d8392b1eaa860eca29d4397d9ac2ae01af00e3505a9d603d5008e7f8599

Observation 0872b2d1-4d96-4fd7-a576-8b93f4c585a0 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:35.367262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:35.367262Z digest=sha256:055ee6add651e80c6d40a38ac04aa8a6bc9dc7909733610453495f74f3f10d21

Observation 4008008e-1e56-4c2a-9d80-8aef88889994 · outbound

This paper cites an unresolved cited work.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction Unresolved cited work

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:33.985256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:33.985256Z digest=sha256:0d623bbaafb8e3c8f6b1eb2b3178438363d410ceca59f320ecc7fe47a4effc96

Observation ac3b1f72-e304-4413-aa12-aa9b707e7c18 · outbound

This paper cites Depthlm: Metric depth from vision language models.arXiv preprint arXiv:2509.25413,.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction Depthlm: Metric depth from vision language models.arXiv preprint arXiv:2509.25413,

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:33.622893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:33.622893Z digest=sha256:019d91a74928a3d36afab863306e397e3082ca576611129d89ea1aa3f68f7ddf

Observation 83e0c2c3-6ba3-4fb7-8596-b705792efcad · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:33.484278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:33.484278Z digest=sha256:efbaf4a8bbf3f3c82c99ef8f9e37836d466802e4e55d507105dd1e1827c7d207

Observation 1e66a50c-6327-46c8-b119-652a6d6c6a69 · outbound

This paper cites UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:34.181055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:34.181055Z digest=sha256:c1b9997a5c167888663346a007e469b834c0a29a519bdcf377df766f72c6a690

Observation 081f0df1-d124-4cc9-abe0-b3032db15574 · outbound

This paper cites SAM 3: Segment Anything with Concepts.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction SAM 3: Segment Anything with Concepts

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:33.793070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:33.793070Z digest=sha256:65cc916f2848d1c22daee1a634259925f145c4e24847772157a42779b7db60a2

Observation c736002b-5762-44f3-92ec-012ce0ce4a4b · outbound

This paper cites Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K., Galley, M., and Gao, J.

DenseMLLM: Standard Multimodal LLMs for Dense Prediction Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K., Galley, M., and Gao, J

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T23:23:34.046942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:23:34.046942Z digest=sha256:75d9826f06c28f79a1318e5b76dcb6a22ab8796e02a8234e10266e735a99ae2c

Pith citing papers

Observation e2fbf0c8-a908-4dbb-a474-6a0934b4ef71 · inbound

Vision as Unified Multimodal Generation cites this paper.

Vision as Unified Multimodal Generation DenseMLLM: Standard Multimodal LLMs for Dense Prediction

Reference 104

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.426147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:e1adc3e78d4ea19dd59f57ff7aef59a8b8f370747560a6f6099e8c8e81086fd4