Pith. sign in

Paper Citation Record · LEDGER

InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2305.05662.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.05662 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:46:49.045303Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:08:55.658612Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dd337dbf-d56d-4135-a28d-da0da5248840 · inbound

VideoChat: Chat-Centric Video Understanding cites this paper.

VideoChat: Chat-Centric Video Understanding InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.598826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:3a4f2b63cd2eabd29e5290461311af77cc328f13b6d1262e1dee2700cdc8389b

Observation feddb060-ac08-41fd-85b3-b92b79efb313 · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.516674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:a0c80df54b20fe008b680f03875ce473b4999d63e793f42489d1b6b12d9f6738

Observation b34065cb-728f-40bc-9c7d-ea7dd6ff454b · inbound

The Rise and Potential of Large Language Model Based Agents: A Survey cites this paper.

The Rise and Potential of Large Language Model Based Agents: A Survey InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 299

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:47:54.533970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-11T10:47:44.152066Z digest=sha256:eed8648143bcd2a7934590fefb1e5c4b04e850e598a931534bda90da503b26d2

Observation ee97f41b-1822-4edc-85b2-da5fcc40a46f · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:46:10.030518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:97c0ab3fd71de7277a41880e800ce4b0404780c0c1e4f8b89c25e1946042823f

Observation 91f3c253-72ec-41f9-855e-271bfc0e9c18 · inbound

Data-Centric Foundation Models in Computational Healthcare: A Survey cites this paper.

Data-Centric Foundation Models in Computational Healthcare: A Survey InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 181

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:13:52.961508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-24T04:13:05.328492Z digest=sha256:f044fba5a1ba1855c6fcb0d18b0fe13ddc75af113d7dbf5b8a22cfdf5d2c71bf

Observation 7336206c-68d3-48c4-b7a7-934d169d5c05 · inbound

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models cites this paper.

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:33:30.353423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T02:33:30.143907Z digest=sha256:b094c85cdef66c845e7c299f00b21d0c51461735d01148a8f9d23642917e007b

Observation 7d452ab8-3858-4dea-a383-87f323cf3352 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.117021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:c92e2aee1ff934b8b6dc543dc482cd43eabba375c8e0643f390673ae2272b482

Observation 6fd60986-4138-4328-9308-0a4cbd49ab73 · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:16:17.302720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:01b8930b53bb0fe47cf35be83fc088ee0eb1ee8badf1a131a2d2c76af35d0530

Observation c226c853-a9fb-4259-98e9-f04c08cf2d81 · inbound

CATP-LLM: Empowering Large Language Models for Cost-Aware Tool Planning cites this paper.

CATP-LLM: Empowering Large Language Models for Cost-Aware Tool Planning InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T13:19:06.905350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:19:06.905350Z digest=sha256:209d7acb7303af198ad34868ff3896aac30faed567d0275fa5a8cf2956d275e3

Observation 8a180739-6da4-4d7b-ba22-f36db095805f · inbound

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model cites this paper.

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T04:22:03.938006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:22:03.938006Z digest=sha256:bf50431f73770801931cbca544d8dfff958f4d672a302d8abaf045439324131d

Observation 37df569a-195a-49a5-bbf3-9c3ae644ea7f · inbound

AIpparel: A Multimodal Foundation Model for Digital Garments cites this paper.

AIpparel: A Multimodal Foundation Model for Digital Garments InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T22:02:03.876613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:02:03.876613Z digest=sha256:0783926c7b10047f0e8de0d9c0ba4130d1ce2bcf48545370265c1046f7aaf177

Observation ddc7b7af-d2c7-45dc-a28a-874ae8c0d7c1 · inbound

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing cites this paper.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.539652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.539652Z digest=sha256:a7307bfe8c92d26728c37ec322d332a40805c86a669e7f97ae2d236da7b5b7f2

Observation 3f7d5a77-2fe2-4f49-b14d-f57a14295374 · inbound

Beyond Turn-taking: Introducing Text-based Overlap into Human-LLM Interactions cites this paper.

Beyond Turn-taking: Introducing Text-based Overlap into Human-LLM Interactions InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T00:42:24.776729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:42:24.776729Z digest=sha256:34638d613f5d7e02ea68d8ef56d8bda6e15d3235ee9337af89d195b5e6f044cb

Observation 53607f26-2daf-445b-b3fa-b9201fa7bfed · inbound

Foundation Model-Driven Framework for Human-Object Interaction Prediction with Segmentation Mask Integration cites this paper.

Foundation Model-Driven Framework for Human-Object Interaction Prediction with Segmentation Mask Integration InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T05:46:49.045303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:46:49.045303Z digest=sha256:c7c30aee6764dca9b951125a5d09daa7f4c692eea8426a7b45cd45b03a12ac41

Observation 9e31acdc-ec99-411c-8518-ebe57bced002 · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:13.117762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:13.117762Z digest=sha256:965d785b32340b927dd78047a845e26c0151a38cba0781f1f9ca3835d35ade63

Observation d21d2863-bf31-45e2-87fa-a91a5f5dbc6a · inbound

Vision Generalist Model: A Survey cites this paper.

Vision Generalist Model: A Survey InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:51.207678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:51.207678Z digest=sha256:3dc45bbe65a68d7ebfccda2357bc6c7ee75cc2819265882e488c802d98a565db

Observation 6e1eb8d8-ba23-4e5a-b01b-fe3be53480d0 · inbound

InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models cites this paper.

InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T18:55:06.722357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:55:06.722357Z digest=sha256:e1c9ad747742caaf63954bd04412dfd8cece85643de99a65a8f8b7255bae0c9a

Observation 4b577851-b868-4a87-a58d-e891afe63e52 · inbound

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance cites this paper.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.773169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.773169Z digest=sha256:01f3e1a9d0f6aabe650f79d9f639c0163c08cce1d439d473e2fd1eef7e7edb2b

Observation 3c10a5ff-4100-4602-8d76-33c8e97745b3 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:02.802194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:02.802194Z digest=sha256:f16d76fc8e9f4e448efbac87229be4bdd449d754865f31bb4ab969236bf20c08

Observation 2e4ab4fc-1b0e-4b94-b490-7cca333b0fa0 · inbound

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning cites this paper.

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T00:59:52.580855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:59:52.580855Z digest=sha256:9f3db0b7ad706df6e8f5b8eef5264e7d47d9295a2dc1ec2e34fc3e762c6345ee

Observation 501209b5-3ea4-4480-985f-e01108c9f3d6 · inbound

Explain Before You Answer: A Survey on Compositional Visual Reasoning cites this paper.

Explain Before You Answer: A Survey on Compositional Visual Reasoning InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 134

Resolution
unresolved
no resolver link, observed 2026-08-15T17:09:18.008871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:09:18.008871Z digest=sha256:972a5ba5042282f0f76729dc980225370e06edba4aa09d94a28eaba7429fcb24

Observation 2f6371ee-f53a-4db8-9243-894cc83ee150 · inbound

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning cites this paper.

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:17.105582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T20:28:30.864143Z digest=sha256:6d46659d6ebceb47ee2428676555cb06de4932520934faf3050313f38a20a31d

Observation 64140c31-6391-473c-8d4b-e3a21d9a9c6d · inbound

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation cites this paper.

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.122437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T09:20:54.635375Z digest=sha256:56d1910881e1ac91c1f934cdd3f89d5ab74e0ff69dc2230695d64098cb663547

Observation 6cac5e5d-667c-4c06-a6b4-08b10b62f271 · inbound

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation cites this paper.

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T16:10:01.893221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:10:01.893221Z digest=sha256:5edb867ad691b49b69b62e52463940e7b09165bc4e8d0b4d721c2b144a23f171

Observation 1529a806-b412-4f95-b41e-e8133415e65a · inbound

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation cites this paper.

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:28:18.704967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T08:04:48.283908Z digest=sha256:06551a64c1e7e102a4b66ebcc556e31fed85bd98c6cee66310cdd07ca43967b1

Observation 2724805f-f09a-4d6c-bea4-08dd5cc7b783 · inbound

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models cites this paper.

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:08:55.660752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T01:42:30.005911Z digest=sha256:d5c2328f7377615b39d15cfc326276c852f8bfae93e38ddf4d6b8ba8b7c5354f