Pith. sign in

Paper Citation Record · LEDGER

NExT-GPT: Any-to-Any Multimodal LLM

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 46 inbound Pith citation observations for arXiv:2309.05519.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.05519 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 46 of 46 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T19:21:22.364061Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T01:19:20.350619Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3df34cfa-faa2-40dd-8ab2-0d43f0cc918a · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.381448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:0ab80f0415a4a0f765501d895f32a8648d97f9e1006fb5db99b1e3addaf78de8

Observation 0a8a8608-cd7e-4852-9e53-25befcb570f2 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks NExT-GPT: Any-to-Any Multimodal LLM

Reference 157

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:46:10.183385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:5777445e65eeb634da1d8aabbff4f9e56f33c407cadae32e833c60c584eff86d

Observation 42510071-7388-47e6-9155-61ae7828ba4b · inbound

Large Language Models: A Survey cites this paper.

Large Language Models: A Survey NExT-GPT: Any-to-Any Multimodal LLM

Reference 218

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:22:56.001147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T15:22:54.023279Z digest=sha256:75fe18e192581f5b14534e5bff2d68ae8c472046dd03b5d97fa1e971d2c6f87e

Observation db948bda-a4a1-4bf3-85e3-84dacd27d500 · inbound

3D-VLA: A 3D Vision-Language-Action Generative World Model cites this paper.

3D-VLA: A 3D Vision-Language-Action Generative World Model NExT-GPT: Any-to-Any Multimodal LLM

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:18:27.288689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T18:18:27.211034Z digest=sha256:17dc7b151fa6aac98ce86db534c0a77b549cb119353e11527ced46cc2556c0b7

Observation 45692325-04a1-46af-8fe8-1c6a85e747e7 · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.326697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:41297498655ead199cbf2ec6e7fab7aa57f7b1ff26050d3658ec8ea643d3cc7b

Observation 2f4d4769-68e1-4ea2-84ef-7833b27cc0eb · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites NExT-GPT: Any-to-Any Multimodal LLM

Reference 124

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.261041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:62aecef89bd43e8589d5c2e768fb78a8b1d5727cea81fca6a5d0da4b263b651c

Observation 082d8ffd-e82a-4384-9e7f-0bc07055dbf8 · inbound

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding cites this paper.

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding NExT-GPT: Any-to-Any Multimodal LLM

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:58:37.460606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T14:58:37.383749Z digest=sha256:e198396e15d4d28a434b94da035f2d8a3a6df068e31a88d34b4bcf983ed16402

Observation b2512e14-5851-4017-9d56-dc3c580554a2 · inbound

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation cites this paper.

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:03:33.624348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T21:03:33.427939Z digest=sha256:7b1651ad8f6a881d76674f9ec63c39a36b3fefd18f7dcb7d95fec7de0113faf5

Observation 20c6e468-7285-4dc5-a580-52853293193b · inbound

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation cites this paper.

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:09:16.322185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T22:09:16.001309Z digest=sha256:6f86f5ef962e9a77054a226b0e5938d3da14c1dfb87dc284cf4dc5e68e4f81b9

Observation 2720e999-b49d-4523-96c1-3ad73b20fae7 · inbound

Analyzing Multimodal Interaction Strategies for LLM-Assisted Manipulation of 3D Scenes cites this paper.

Analyzing Multimodal Interaction Strategies for LLM-Assisted Manipulation of 3D Scenes NExT-GPT: Any-to-Any Multimodal LLM

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:43:18.972547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T18:42:00.624044Z digest=sha256:854792b9638cc5166f296b68f996d87e095eadc7c4dcb6a1ec05a966b09a7e0c

Observation 5f43f488-8b82-49a8-898b-b42f95dcb212 · inbound

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling cites this paper.

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling NExT-GPT: Any-to-Any Multimodal LLM

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:14:53.004260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T08:14:52.890145Z digest=sha256:96f38e3541deffcf8609067b2737eba81db48d321b2700b2ffabc5f6b6292349

Observation 871c4183-e5c8-4bfa-98bc-e6d6e3c02401 · inbound

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning cites this paper.

NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning NExT-GPT: Any-to-Any Multimodal LLM

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:22.364061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:22.364061Z digest=sha256:f088ea13a5f5dc27c01e923a4b42b36b8a789eb825210e4d79f90ca2946de552

Observation eec755af-05c5-4871-b90c-f65d67021776 · inbound

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding cites this paper.

UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding NExT-GPT: Any-to-Any Multimodal LLM

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T19:32:32.021100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:32:32.021100Z digest=sha256:b603fbbfd6084c6c0e57adc939bb94e65fd07f51188d2bbbe2c87ddf3db17e1d

Observation cb9d029c-c0d8-463e-80cb-d7253de4152e · inbound

UniCoRN: Unified Commented Retrieval Network with LMMs cites this paper.

UniCoRN: Unified Commented Retrieval Network with LMMs NExT-GPT: Any-to-Any Multimodal LLM

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-08T05:54:16.104506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:54:16.104506Z digest=sha256:6c39b72a2f28f0fe5e66711e32447e483945d95fe8db40d7d0840015c8449f8a

Observation 02be9b8b-2e7d-40e8-9c38-3db5ebe977c9 · inbound

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation cites this paper.

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation NExT-GPT: Any-to-Any Multimodal LLM

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:30.561275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:21:30.561275Z digest=sha256:06997c20ddf180196f10f1e66930c10d5ae132dccced18c0a9ba9af8049ffd3c

Observation adc24064-c836-4590-849d-ead80d009c80 · inbound

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization cites this paper.

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization NExT-GPT: Any-to-Any Multimodal LLM

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:04:22.755345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T15:04:22.690503Z digest=sha256:016e37a04e3138e19b6a7a8a6ffe57b301b87c130beeb3d451a7c3a5c7ff3107

Observation 92929f70-6113-4020-9115-e3a98e64831e · inbound

Transfer between Modalities with MetaQueries cites this paper.

Transfer between Modalities with MetaQueries NExT-GPT: Any-to-Any Multimodal LLM

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:49:23.284023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T22:49:23.074271Z digest=sha256:9ec391242cce065e96a48b8facd87b1f5b909521dfe679304c6b0a2efcb229e8

Observation a07c77f3-61e7-4fa2-9595-d80b5e5c4b9d · inbound

Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method cites this paper.

Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method NExT-GPT: Any-to-Any Multimodal LLM

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:42.109858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:42.109858Z digest=sha256:7faac118c4360428e378036e7f4dbf990464663eea8126d4369d7cb5d28f8514

Observation 489727a3-afcd-4908-9d9c-f9f8f5e0b8df · inbound

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models cites this paper.

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:03.563984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:23:03.563984Z digest=sha256:ecd3727f45fe1a1a47a3716acf884908f6955e7a745c6b4fdfef0911bbaed62e

Observation bfb03666-dc73-483e-a948-ce00f2e52775 · inbound

MMaDA: Multimodal Large Diffusion Language Models cites this paper.

MMaDA: Multimodal Large Diffusion Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:59.768622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:2476cf3f6a214db71a299e22f102e89c80eb4c3bc143601d3a3c3bc3ad60c1d3

Observation 1e3af57e-dd61-4774-a7a4-6d5ab4db2843 · inbound

Sensorimotor Self-Recognition in Multimodal Large Language Model-Driven Robots cites this paper.

Sensorimotor Self-Recognition in Multimodal Large Language Model-Driven Robots NExT-GPT: Any-to-Any Multimodal LLM

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:32:19.269257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T13:29:34.151546Z digest=sha256:9a168b88ea782f5a4943ba1e29e125826a7efa296b3cb8fa0456c676aade08b5

Observation f9832be2-b9f5-4fdd-aeac-e089fbae9e01 · inbound

Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion cites this paper.

Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion NExT-GPT: Any-to-Any Multimodal LLM

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:06:49.990929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:06:49.990929Z digest=sha256:ae41b9b4b9b4d40cf704c6db3a5951c983f531727bb3b1734ef8ed4b5353f1de

Observation 727c3c1d-52ce-4d27-bf43-8b46fbc2fb7f · inbound

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning cites this paper.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning NExT-GPT: Any-to-Any Multimodal LLM

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:26.460651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:26.460651Z digest=sha256:cbd5f0efb5fb2e4feb0758d96ffe636b07cf51676abd6331772550ef338e381e

Observation 461ee106-6d4a-4934-8cdb-427e0f7fba8f · inbound

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation cites this paper.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:20.351944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:20.351944Z digest=sha256:59c5ac430a3a8c0e25cc1e6affe25e833b54efe7f20cc3f2443c86fc7176d169

Observation 7684d983-6c48-4c85-9bb7-c98a34ce78cb · inbound

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems cites this paper.

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems NExT-GPT: Any-to-Any Multimodal LLM

Reference 197

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:52:16.135089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T11:49:36.574471Z digest=sha256:b661ad45572820931f8ac45726be441798c7b06c6c557dc3ae7bcb4c07defc95

Observation 48c745e6-6456-4854-beba-54ef90eaad78 · inbound

Multimodal Representation Alignment for Cross-modal Information Retrieval cites this paper.

Multimodal Representation Alignment for Cross-modal Information Retrieval NExT-GPT: Any-to-Any Multimodal LLM

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:08:32.375241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:08:32.375241Z digest=sha256:785746fce4bb5a55b3731b36c2443d6f6a9449dc2c239964fca6e1bac8b0d09c

Observation d18c754e-cc81-41f7-95a0-bf11ed74dfb9 · inbound

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation cites this paper.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.517877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:09.517877Z digest=sha256:336ff2bf762ebecff3ac79cb713700afe95262c068843e660bd986f7da521b5e

Observation bfead915-d419-4564-9a55-6c6beaa311aa · inbound

DanceChat: Large Language Model-Guided Music-to-Dance Generation cites this paper.

DanceChat: Large Language Model-Guided Music-to-Dance Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:26:42.724376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:26:42.724376Z digest=sha256:7635214e194ec58ea22dcc0120496d31ed1ee4c094f4c191b45017ea61dd92f1

Observation 23ba9004-94bb-42f2-8730-5eecc4029b98 · inbound

Enhancing Hyperbole and Metaphor Detection with Their Bidirectional Dynamic Interaction and Emotion Knowledge cites this paper.

Enhancing Hyperbole and Metaphor Detection with Their Bidirectional Dynamic Interaction and Emotion Knowledge NExT-GPT: Any-to-Any Multimodal LLM

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:59:09.376377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:59:09.376377Z digest=sha256:3e9e7fbca98bdf637b3d50725072c070c8f4d4127c14a171d84a03e840b6b943

Observation 3dc49e84-f5b8-4392-8dd6-9ffc27b499d3 · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 120

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.793866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:1f9f020c82c8e1ccd5f813f1075e750b8197342db158d222ce2be083e71d6141

Observation 9950c082-bd59-48a2-b4f5-f46c7fa7a7a6 · inbound

NeoBabel: A Multilingual Open Tower for Visual Generation cites this paper.

NeoBabel: A Multilingual Open Tower for Visual Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.821240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.821240Z digest=sha256:8073de925f341f07e8d35f76f62a7130e5214afd5bd6ae2a397ccb03a970cb3c

Observation d056821a-5da2-4023-a2fa-faf90712d6b6 · inbound

Learning Deliberately, Acting Intuitively: Unlocking Test-Time Reasoning in Multimodal LLMs cites this paper.

Learning Deliberately, Acting Intuitively: Unlocking Test-Time Reasoning in Multimodal LLMs NExT-GPT: Any-to-Any Multimodal LLM

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:55:33.904321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:55:33.904321Z digest=sha256:fe96b3bbb33aeae30492300972cca49a7ae6b7002f7f9c687741df189f98a809

Observation 98fc655e-ce05-4da9-9151-2a93f7c08413 · inbound

DatasetAgent: A Novel Multi-Agent System for Auto-Constructing Datasets from Real-World Images cites this paper.

DatasetAgent: A Novel Multi-Agent System for Auto-Constructing Datasets from Real-World Images NExT-GPT: Any-to-Any Multimodal LLM

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:59.162850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:59.162850Z digest=sha256:6a03248b14b26c8013a0bcb585fbac1e344615c6617cc183a26bbdb452d32e4a

Observation 9d6818eb-e463-4cf5-8a03-e44237d8a785 · inbound

Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan cites this paper.

Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan NExT-GPT: Any-to-Any Multimodal LLM

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T05:49:51.608775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:49:51.608775Z digest=sha256:8a0452aee0a253926585f3da719f7aef41365d5ed0ed9307b1ea801f4f6ecf35

Observation 0d76b3bd-d2d5-474b-bdc9-8f2c483b45cd · inbound

Audio-Guided Visual Editing with Complex Multi-Modal Prompts cites this paper.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts NExT-GPT: Any-to-Any Multimodal LLM

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.833708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.833708Z digest=sha256:98fe7dd9c775e1a86edca384aec48ec3d80c7b53283a67aae064500fa94ee577

Observation 85d0e4a6-b2b6-458c-a997-529ec7b49f2e · inbound

Effectively obtaining acoustic, visual and textual data from videos cites this paper.

Effectively obtaining acoustic, visual and textual data from videos NExT-GPT: Any-to-Any Multimodal LLM

Reference 121

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:37.086887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:37.086887Z digest=sha256:1191530bf436b13b3cadcd27de3e7fc26ac9cde6278a3aa68bfb74ca2de23e21

Observation b51635ce-46ee-4444-b6c1-0933ccef8253 · inbound

Testing chatbots on the creation of encoders for audio conditioned image generation cites this paper.

Testing chatbots on the creation of encoders for audio conditioned image generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.343930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.343930Z digest=sha256:c280da5d781856373eb4d2410380e448b5e62224e248bd44ed55536344d07611

Observation d8e3f15a-3507-4bad-b7e8-741d58f22007 · inbound

Cross-Modal Backdoors in Multimodal Large Language Models cites this paper.

Cross-Modal Backdoors in Multimodal Large Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:20:55.597232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:51:11.432424Z digest=sha256:6d72f54cb0cb606f42bc5796942b1486512939436d47c2c4e479df1a5abe2078

Observation 5ed9a3ef-3f09-45c9-ac65-4cb9505567fa · inbound

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction cites this paper.

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction NExT-GPT: Any-to-Any Multimodal LLM

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:38:19.150954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T13:36:44.071188Z digest=sha256:d81c031b6b8c4dbec1cef86c39f15bc40a66fe23a5ad5fb430a07c9722952f48

Observation cce32fc9-da56-4075-a383-346ea8e357f4 · inbound

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction cites this paper.

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction NExT-GPT: Any-to-Any Multimodal LLM

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:19:20.353196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-04T01:11:42.073993Z digest=sha256:42c46cb3631504a2d8c020963ffec01f033f093dcfc49db079448aefbb3fa8b0

Observation 447a55f8-d548-43a5-b154-53f0f8606604 · inbound

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation cites this paper.

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:28:18.747749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T08:04:48.283908Z digest=sha256:677e90a7e7d9a1036cd26b5e7f107dd781d0793ec3ecf95a0d1ba745a651fc26

Observation c54697af-3353-492b-b83d-ca682f3beb8a · inbound

EmoteGPT: 3D Human Facial Expressions from Natural Language Descriptions cites this paper.

EmoteGPT: 3D Human Facial Expressions from Natural Language Descriptions NExT-GPT: Any-to-Any Multimodal LLM

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-12T07:49:23.168452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:49:23.168452Z digest=sha256:d915c507ca9950ac294dc35dccb2b63b6fb69919de0e3ca2e32b49f1eef79568

Observation 984aec55-eda8-4a39-a2e3-559eded843aa · inbound

Laguerre Geometry for Interpreting Large Language Models cites this paper.

Laguerre Geometry for Interpreting Large Language Models NExT-GPT: Any-to-Any Multimodal LLM

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T10:42:30.055074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:42:30.055074Z digest=sha256:a08e79ea0cad05a9c09e83768618a784d037a98ef7a4c6942e6c0f5336888270

Observation 0f404c44-6f12-46ba-bceb-19994e9b0001 · inbound

S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation cites this paper.

S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:27.351494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:38:27.351494Z digest=sha256:8ccbf8f255fdc30525618fdd5062db216cf1f8bc2101604fa3a89da491ac2933

Observation d955bed7-164c-439e-ba02-5d2da90d9fc5 · inbound

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens cites this paper.

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens NExT-GPT: Any-to-Any Multimodal LLM

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-03T08:35:55.978209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:35:55.978209Z digest=sha256:a5744e6004e27bd7f454f627caf46da8515db2575ece48678290deba96f311e3

Observation 3c2af9b0-811a-4563-83a0-7961af6d3030 · inbound

GARDRec: Decision-Level Graph Grounding for Large Language Model Recommendation cites this paper.

GARDRec: Decision-Level Graph Grounding for Large Language Model Recommendation NExT-GPT: Any-to-Any Multimodal LLM

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T00:55:47.192966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:55:47.192966Z digest=sha256:1f56ef4a16649c42662ec016ce8597690eb2d492407b86acb2937762c9708e40