Pith. sign in

Paper Citation Record · LEDGER

InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2305.05662.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.05662 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:43:51.207678Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:08:55.658612Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dd337dbf-d56d-4135-a28d-da0da5248840 · inbound

VideoChat: Chat-Centric Video Understanding cites this paper.

VideoChat: Chat-Centric Video Understanding InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.598826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:0f59ba6f9d9c2278ad7d5df1be1f318a813d63e8e89512f97a2d09084fbe3ebc

Observation feddb060-ac08-41fd-85b3-b92b79efb313 · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.516674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:f0cb1e9e500d466e3df9dd06495aa31e83bf20174d7e73f1589012a87fda39dc

Observation b34065cb-728f-40bc-9c7d-ea7dd6ff454b · inbound

The Rise and Potential of Large Language Model Based Agents: A Survey cites this paper.

The Rise and Potential of Large Language Model Based Agents: A Survey InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 299

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:47:54.533970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T10:47:44.152066Z digest=sha256:c9c200c06580e179ff9653fe936e7480440f5a785411e25d66695db6939f6ec4

Observation ee97f41b-1822-4edc-85b2-da5fcc40a46f · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:46:10.030518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:55c89f3f65fdd3a269bc80a93a046761fb1851554fe6adb27ad6f138e8226adf

Observation 91f3c253-72ec-41f9-855e-271bfc0e9c18 · inbound

Data-Centric Foundation Models in Computational Healthcare: A Survey cites this paper.

Data-Centric Foundation Models in Computational Healthcare: A Survey InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 181

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:13:52.961508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T04:13:05.328492Z digest=sha256:6e5fe3507f00a759d55d4dce8f958026b1e65f15c750d71c1c172f5122a52ff1

Observation 7336206c-68d3-48c4-b7a7-934d169d5c05 · inbound

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models cites this paper.

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:33:30.353423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T02:33:30.143907Z digest=sha256:b5324cb0b7e7b6d85145a84fd4655fc92928fde396348acc2212b57a9fdb9635

Observation 7d452ab8-3858-4dea-a383-87f323cf3352 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.117021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:4f6cb496477c680ce21d33b6e439c0f77f6615f9e95189202a5f2667b276574b

Observation 6fd60986-4138-4328-9308-0a4cbd49ab73 · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:16:17.302720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:24d35e814d8430e29f9c196ab422184509c7408e6efbd19e9e652522fbfde944

Observation d21d2863-bf31-45e2-87fa-a91a5f5dbc6a · inbound

Vision Generalist Model: A Survey cites this paper.

Vision Generalist Model: A Survey InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:51.207678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:51.207678Z digest=sha256:87af394903f886b4dfaaf0daa51a5a00307e75f18a697c5c819b8e677df61a47

Observation 4b577851-b868-4a87-a58d-e891afe63e52 · inbound

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance cites this paper.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.773169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.773169Z digest=sha256:c7b0baea25c40b3561bbcc41b4d4b135a07b84c87b565d69c273bb612619836a

Observation 3c10a5ff-4100-4602-8d76-33c8e97745b3 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:02.802194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:02.802194Z digest=sha256:0efba1d6054f54d5cb9c2a5f6b0fb40eb36bbf95e2bd416410266d56c1040b40

Observation 2e4ab4fc-1b0e-4b94-b490-7cca333b0fa0 · inbound

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning cites this paper.

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T00:59:52.580855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:59:52.580855Z digest=sha256:8942bf10dc9db24af3e4c8963bcd0ce4eb7482342d94b4f8b9be347ad91200b5

Observation 2f6371ee-f53a-4db8-9243-894cc83ee150 · inbound

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning cites this paper.

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:17.105582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T20:28:30.864143Z digest=sha256:49185434985e4307cbe748544662aae10ce94c5b454b104e0b2506c40c1710f1

Observation 64140c31-6391-473c-8d4b-e3a21d9a9c6d · inbound

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation cites this paper.

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.122437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T09:20:54.635375Z digest=sha256:94c18973af2cac9789ccd6299664189cabb060d3ae1238f1fe13f9b132c554b3

Observation 6cac5e5d-667c-4c06-a6b4-08b10b62f271 · inbound

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation cites this paper.

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T16:10:01.893221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:10:01.893221Z digest=sha256:e652cc1b117c2c09566799779edb971829ea2397c64e3c8ad5dda848df97fd58

Observation 1529a806-b412-4f95-b41e-e8133415e65a · inbound

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation cites this paper.

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:28:18.704967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T08:04:48.283908Z digest=sha256:39a4b4bacec35a98c3f965b52d1a2380a64e8838e93c38b3f58c74578bcc17dd

Observation 2724805f-f09a-4d6c-bea4-08dd5cc7b783 · inbound

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models cites this paper.

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:08:55.660752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T01:42:30.005911Z digest=sha256:350b89cfe36b18d9181aa02e6464edec6a4e9187c4fd1afab927afb8b35b8aa3