Pith. sign in

Paper Citation Record · LEDGER

NExT-Chat: An LMM for Chat, Detection and Segmentation

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2311.04498.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.04498 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:47:12.501355Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.195600Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 905fa20d-3f0a-40a2-a892-68407df11eb5 · inbound

MiniCPM-V: A GPT-4V Level MLLM on Your Phone cites this paper.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone NExT-Chat: An LMM for Chat, Detection and Segmentation

Reference 116

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:32.136962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:23479ccb73e561b7c3a115cb459d1f9847a3dfa48caeb1ce9d07a2003cc74fc2

Observation b66f6c78-d26d-44b3-95be-f2a4daf22cb6 · inbound

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment cites this paper.

Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment NExT-Chat: An LMM for Chat, Detection and Segmentation

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-11T00:47:12.501355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:47:12.501355Z digest=sha256:ab5f7e886d2b5dc8a3c1e871e921e60b929dcdb37baa43264c499893cffeb213

Observation b6a900c4-f51b-41f0-b3c0-e35bfa0503ab · inbound

Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging cites this paper.

Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging NExT-Chat: An LMM for Chat, Detection and Segmentation

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-10T23:40:45.252649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:40:45.252649Z digest=sha256:f4600472905687981368f4c93c470279c1d271fab8953021d0f0edc65f295f51

Observation 36949f8a-0830-4cdb-9115-913f2a23aded · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications NExT-Chat: An LMM for Chat, Detection and Segmentation

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.252943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.252943Z digest=sha256:15882d2afba8af9d9f5749297ba94f361b41f3c3f37aa8dfa7236eb461883d8b

Observation 3470eee5-4f32-451f-96d3-38bd064a9e36 · inbound

CityLoc: 6DoF Pose Distributional Localization for Text Descriptions in Large-Scale Scenes with Gaussian Representation cites this paper.

CityLoc: 6DoF Pose Distributional Localization for Text Descriptions in Large-Scale Scenes with Gaussian Representation NExT-Chat: An LMM for Chat, Detection and Segmentation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:16:33.505694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:16:33.505694Z digest=sha256:827b2d663a7c24bcd5a3d43bc6fe8d87ba967438c9e31e50d15b927ba2ccde6e

Observation 64f6bba2-b9e5-4be4-a078-cbdb17883279 · inbound

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling cites this paper.

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling NExT-Chat: An LMM for Chat, Detection and Segmentation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:52:20.818808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T02:52:20.643070Z digest=sha256:3ed3aaf99326de273a2501bb778583095450604799405a84525558b0dd2d701b

Observation 8df77128-ee58-44d8-88f3-598d45721f14 · inbound

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation cites this paper.

HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation NExT-Chat: An LMM for Chat, Detection and Segmentation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T16:43:56.096102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:43:56.096102Z digest=sha256:67acb8c18830bf38d30c03fbecac26edac2a34403567d9048021362c7693b55d

Observation ee2978c9-438b-4601-bb5e-388ffae6413d · inbound

ReMeREC: Relation-aware and Multi-entity Referring Expression Comprehension cites this paper.

ReMeREC: Relation-aware and Multi-entity Referring Expression Comprehension NExT-Chat: An LMM for Chat, Detection and Segmentation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T15:16:25.400535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:16:25.400535Z digest=sha256:b1ff29af279748cdb9ffe311f4a10b058f9bce4cb21ce15af5d129d11d8fe745

Observation 0b5d19e3-4542-4384-b022-8553bc7d6b45 · inbound

MediRound: Multi-Round Entity-Level Reasoning Segmentation in Medical Images cites this paper.

MediRound: Multi-Round Entity-Level Reasoning Segmentation in Medical Images NExT-Chat: An LMM for Chat, Detection and Segmentation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T22:08:38.655224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:08:38.655224Z digest=sha256:e7f95dcf7942d90e19ecfa12f9d86003b734e658073ee59005f2d056dca441af

Observation a988e8a7-e882-4a9d-969d-d4f191512d13 · inbound

Grounding Everything in Tokens for Multimodal Large Language Models cites this paper.

Grounding Everything in Tokens for Multimodal Large Language Models NExT-Chat: An LMM for Chat, Detection and Segmentation

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:31:21.914672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T23:31:05.422935Z digest=sha256:77d41a248a43029570f384a10d583213771bae31213b02c7889ce40ce74896dd

Observation 00f6c03f-552e-4c09-9edf-3245116f0303 · inbound

RemoteAgent: Bridging Vague Human Intents and Earth Observation with RL-based Agentic MLLMs cites this paper.

RemoteAgent: Bridging Vague Human Intents and Earth Observation with RL-based Agentic MLLMs NExT-Chat: An LMM for Chat, Detection and Segmentation

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:40:59.609219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:00:20.216268Z digest=sha256:40c178caa3f3bc9e67316cf828efc0306dafb891b152061f3fa1602d564805c0

Observation 1adf90e1-9e2c-4522-a5a5-ff1336077c44 · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation NExT-Chat: An LMM for Chat, Detection and Segmentation

Reference 223

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:08.095451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:b5e6817e740b755d48c4ec50ef048b68a1bb28e66d92b690a9499b9601bdc868

Observation 7389e0fe-44dd-4cf4-962e-228af96f60a8 · inbound

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation cites this paper.

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation NExT-Chat: An LMM for Chat, Detection and Segmentation

Reference 196

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:43:49.462541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T05:10:44.608959Z digest=sha256:ab5e13338d2777bc2552447092be1e886d13f3286031b00777b2439855313bca

Observation 94e2fd50-8d09-4487-a8e6-b171afce61fb · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning NExT-Chat: An LMM for Chat, Detection and Segmentation

Reference 104

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.076990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:56acdd152e8985fdf9de6b2fd9bef4a117df5696bd019bf96693f8d26e7e65a4

Observation f3c278d6-bd05-46b6-9d72-b4fdd40d153a · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models NExT-Chat: An LMM for Chat, Detection and Segmentation

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.197378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:b664431bddf59f027ef0dbf88bbc66b9d4d067972c44057076dfa0c64e0aba15

Observation 210014b3-c990-4050-8b71-41f38297f8e5 · inbound

Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose? cites this paper.

Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose? NExT-Chat: An LMM for Chat, Detection and Segmentation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T10:20:57.384780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T10:20:57.384780Z digest=sha256:431f61cc4aa35537d2b211ca3b83a5e8281d0a72527c987df03064990f37d34c