Pith. sign in

Paper Citation Record · LEDGER

AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2402.12226.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.12226 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:53:24.189799Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T00:27:50.594729Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 19aef4f9-61fe-45d6-a1c1-b9b5cf7cd034 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 149

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:41.968189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:17638a12b1a00a692cd0024dd36fd63fad1164c07b187bdabd906468c07b1bba

Observation d3f4e443-b1cd-43da-8290-b36562c5ec4a · inbound

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models cites this paper.

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:44:47.467568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T07:44:47.355960Z digest=sha256:54f142a2264c264d3389fd251fef64588ea98842b1abd6e68f9d17a804c80ce8

Observation a583e6f5-cf07-4c80-bf93-2be4a2b2c17a · inbound

VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation cites this paper.

VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-16T00:26:21.388246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T00:26:21.313005Z digest=sha256:676fe87aaaebe7bd0767f8bc7f6d79b887d6f10ea80a2697bede845ab79c9beb

Observation b198f446-cef7-4eb5-8110-8581f14aece3 · inbound

Deep Multimodal Learning with Missing Modality: A Survey cites this paper.

Deep Multimodal Learning with Missing Modality: A Survey AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:26:03.959221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T22:26:03.706725Z digest=sha256:5913a938aedd4f47dd1e6fa054c4e5d0198eb47922ec32e61c6c8484274b5e63

Observation 952a0763-c33e-4044-a508-360593a97f5d · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:08:19.612711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:398e02d84440b32d3c49ef35c98d51d24fb2eb61b1491f5b0680ae216e911e26

Observation f1209626-3438-41a3-9db3-3db1e996524a · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 213

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.088051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:dd8ef2f3de195f113d2dc49aae0e4067ffbae8fc542084cd72ba31dfae784d92

Observation c326b9e8-d853-4f1f-bc4b-c562c5c40dfb · inbound

Qwen2.5-Omni Technical Report cites this paper.

Qwen2.5-Omni Technical Report AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-10T17:54:03.371019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:54:03.225439Z digest=sha256:cf0d3a80c526e6a76f99183f564dbf1d0157db7643a9deca2c05a63fd1ee26af

Observation e8c6534f-ef50-4d68-aef2-43dc18d60de8 · inbound

SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings cites this paper.

SpeechLLM: Unified Speech and Language Model for Enhanced Multi-Task Understanding in Low Resource Settings AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T13:53:24.189799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:53:24.189799Z digest=sha256:dda466ec000157ddeebc86947a6dd7e2f3a4beaa69971e3b85298ce0e18faf0c

Observation 865e7739-98d2-4c3d-b428-2484a2811029 · inbound

Two-Dimensional Quantization for Geometry-Aware Audio Coding cites this paper.

Two-Dimensional Quantization for Geometry-Aware Audio Coding AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:20:29.187851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-21T18:16:51.486807Z digest=sha256:39080e3a75af38072c58d35ac04f6e3c52ae86ea9a544be9781e49cbcd74b388

Observation 00eb7ac9-d2fb-49b4-8f37-42ed96fc2aed · inbound

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body cites this paper.

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 130

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:08:36.216907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T22:04:07.403410Z digest=sha256:fa69e7b513d805a0ad892ecc7b34eb63d533f222e16d3650069375b6410198c7

Observation 46f89378-7081-4658-9560-11410ee6f9cb · inbound

Benchmarking and Enhancing VLM for Compressed Image Understanding cites this paper.

Benchmarking and Enhancing VLM for Compressed Image Understanding AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:26:41.821712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T07:26:06.192538Z digest=sha256:acecc60b064eddba005a908ad6d426d20f4c80e45609400f360f5b9751108b20

Observation 9394b685-dcd2-4ef3-8b50-a2ee92c9b7b2 · inbound

PolySLGen: Online Multimodal Speaking-Listening Reaction Generation in Polyadic Interaction cites this paper.

PolySLGen: Online Multimodal Speaking-Listening Reaction Generation in Polyadic Interaction AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:30:54.803391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:28:26.910105Z digest=sha256:6650c51ee9971d6438cd8fdefa0fed5482309bb0efb0247351867216244bb5de

Observation df7444ad-9e6d-4463-ae30-47a181b20c57 · inbound

Context Unrolling in Omni Models cites this paper.

Context Unrolling in Omni Models AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:04.879271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T22:02:57.841111Z digest=sha256:d85701debc5f8a7fea3eea7a93e8599ab804ff9227fd404591f5723bd1e6f0ce

Observation 7262370a-9725-46c9-8304-e10518785462 · inbound

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs cites this paper.

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:02:06.312638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:00:02.786195Z digest=sha256:0549b8dbed09dd6387679b834a428ca261ce9d1e24b16a7a68aa7a58d857ebee

Observation 5de52724-6a68-48e7-ae09-d24d0e158630 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:01.631420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:b53e0499a196a2fb954e3be62987e3f5fa3b26827b907703af5d140bcf45c664

Observation 01482320-8175-4f71-ae28-60370b3353d9 · inbound

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs cites this paper.

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.398632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T14:36:53.295540Z digest=sha256:7a7a6572b34e7c40c37b0136ec722ac6b846958c05efcb68f5b8ec38d0afbf9f

Observation 446a62ad-e5fb-4165-93e0-617d8e6b943a · inbound

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment cites this paper.

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:08:37.416070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T06:05:26.735340Z digest=sha256:2d539b7562d4c818f47f19c7c367757aa418d80c870ae96dacae0405105805fd

Observation d7115c61-2317-449d-9185-b06cd9508c8e · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 300

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:44.792436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T09:19:50.623741Z digest=sha256:2eeb2b1bebb596ffdec400322e09d6a370b9c62999e909c6ad6a6e0872ec23e8

Observation e2588dcf-7472-4df6-89d4-15c112287ea1 · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 299

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:55:59.424290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T01:18:19.195007Z digest=sha256:28ce644d2001bf3e0ad43f6ae9689226ed9c8fa40af95165597f7d7157cfdb70

Observation 4021b98c-f757-48e0-8b11-34f352a0c298 · inbound

POPS: Recovering Unlearned Multi-Modality Knowledge in MLLMs with Prompt-Optimized Parameter Shaking cites this paper.

POPS: Recovering Unlearned Multi-Modality Knowledge in MLLMs with Prompt-Optimized Parameter Shaking AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-11T00:27:50.614075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T00:23:30.377235Z digest=sha256:85e95d955bd13dbf35675ff7d3cebb26de9105748102282319bc594bfe005eec