Pith. sign in

Paper Citation Record · LEDGER

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation

As of 7 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 2 inbound Pith citation observations for arXiv:2506.10395.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10395 v2

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:34:12.154340Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T19:17:19.634242Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T18:51:15.880853Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact2
  • verified fuzzy18
  • unresolved65
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eebf5b90-1bdb-45b1-a9e5-af521b65bd51 · outbound

This paper cites write newline.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:59.234925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:59.234925Z digest=sha256:1bdb61d85dabc4eac3b530c0ced335f955f6331fcf37fa7ffb2e5956dd440b4f

Observation a020a70d-a6a7-4c98-9182-37b551e521db · outbound

This paper cites The Llama 3 Herd of Models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation The Llama 3 Herd of Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:59.321702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:59.321702Z digest=sha256:c73c239801aac7cadc14281dedc95e5cf5f8dc3ffb563c72f51e39c4713ada4e

Observation abc4ea03-17c9-4738-936f-47ed169a939d · outbound

This paper cites CM3: A Causal Masked Multimodal Model of the Internet.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation CM3: A Causal Masked Multimodal Model of the Internet

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:59.440871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:59.440871Z digest=sha256:bcceffd31955f987cded378e7fc2cf490f1693d6ef3655bfce97052ab0986992

Observation 28a22b10-b9cf-4c6f-9bdd-35e81f4db077 · outbound

This paper cites Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar \' e n Simonyan.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar \' e n Simonyan

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:59.594642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:59.594642Z digest=sha256:10c14ff9c5840c6ab7e76d8ea022701d6e8c6e7bea2de32f3b38a64c5b87ac06

Observation 9c6ba7fd-8ebb-4380-a453-fdc378c12828 · outbound

This paper cites Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:59.735365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:59.735365Z digest=sha256:fd055cebb918dfe615bbc1de9cd716b222bc95dbfbb02facb05992c0eff9b545

Observation 85ca09f3-7f82-4bec-868e-88e0543cf97a · outbound

This paper cites Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:00.049480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:00.049480Z digest=sha256:7860c76921f3a4c2b0ed5971370320d8085df79cb51b518da8cd1bb6a9eaef33

Observation 0f708b75-a8e2-4530-91a4-596384414343 · outbound

This paper cites Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:00.161811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:00.161811Z digest=sha256:993bb844e6614c8eb83f1b212fcde8d8cf9ed5cee7640a75a506f32dbd2fec2f

Observation 2af2ac08-d9d3-4a7e-823b-559c41fbd359 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:00.326020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:00.326020Z digest=sha256:ec7ee0b2efc6c8fabd176278a971d1ab3adeba0e47ce63769238450651f9cebe

Observation 810f32d4-c9b2-4610-8376-0015a635568e · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:00.478226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:00.478226Z digest=sha256:33780b2e30b46fb45cd51f96bbc8970e800b25bde2693e6d7ac9c348d7ee1352

Observation fdd57776-9a51-47f4-9f03-ec46aa1e85e7 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:00.637220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:00.637220Z digest=sha256:3a223ed4576b47cbd504aed64cd6ba2375e70d5e983960bd9feedb5b4aa18fc1

Observation 282b3c3f-0bf0-434e-a7b2-4d2c58bc8d31 · outbound

This paper cites Dream LLM : Synergistic multimodal comprehension and creation.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Dream LLM : Synergistic multimodal comprehension and creation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:00.829274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:00.829274Z digest=sha256:fe4907038eac089f655091e91454278b2a185252ae93184a4b868ce7ad40b289

Observation 31c365d3-95ab-449f-a0bc-836c440bcaa2 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Taming transformers for high-resolution image synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:00.945047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:00.945047Z digest=sha256:ee5eb41c359aa26ddfd01e38dd08dfef3eba0eb1a28ac9887da7e006cff3432a

Observation 0a84b836-d5e9-4f8a-b917-d0214a4a70ce · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Scaling rectified flow transformers for high-resolution image synthesis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:18.805912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:34:01.070818Z digest=sha256:83b038a4a60f733be05aef053149122386f97613407f4ee8054ca21066b7e02f

Observation 005b0f3d-381c-4faa-8df6-bd1e16f6c3e6 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:01.269001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:01.269001Z digest=sha256:135b3bb04657f3f2ef7a2d97d979fdffecfa3c43c2fc5c0a2065b7192a325a1d

Observation dd691fa1-fb58-40bf-b061-aef520ac4b56 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:01.423591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:01.423591Z digest=sha256:040da3f16ed08e1b602ba9bad8de2e16886e0e60320efa9c553dd3974246b23a

Observation 176c3921-7018-4cdd-8b03-1bf1c652af5b · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:18.610344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:34:01.595680Z digest=sha256:0bba0baca9e8fb27c4eaa5cb961db06451ed4e79ab7d8808d91ed9e558c1c764

Observation 8d0a0171-406d-4d70-be15-7df991db1f82 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:18.408391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:34:01.784116Z digest=sha256:5654ab6f1c1015a9da804d08cec79c5c2114e2333fa78281135c093b0b09061e

Observation 4cc5fe1e-282d-46d5-9df9-8d11ba897449 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Vizwiz grand challenge: Answering visual questions from blind people

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:18.190173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:34:01.983831Z digest=sha256:53dfd600a4a51929081e03446bb24254bf5f7def14adade496042edccc1c891c

Observation 3ecddb5c-dceb-48f4-9232-86434787a8e4 · outbound

This paper cites Masked Autoencoders Are Scalable Vision Learners.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Masked Autoencoders Are Scalable Vision Learners

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:02.155166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:02.155166Z digest=sha256:fd140f96059c4d800e3ee54cb7e7e20467c737d5b7bb01004863eedf02755295

Observation 29f63196-c174-40ca-bc2c-c11ce4af7382 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:17.952933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:34:02.302887Z digest=sha256:f852107246595ddfcf59a4c6e13ac342aa854ce2b00433550439e371a1c91d91

Observation c7a3e172-d7a5-4053-9f1c-a43cb76fc897 · outbound

This paper cites Unified language-vision pretraining in LLM with dynamic discrete visual tokenization.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Unified language-vision pretraining in LLM with dynamic discrete visual tokenization

Reference 22

Resolution
verified exact
doi, observed 2026-08-07T04:34:13.304814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:34:02.418863Z digest=sha256:303e06bff759c59246dac934038f40ec7d22124465bed0027dc67f6fd81f0f2e

Observation 61a80a70-4b79-4239-9a52-ac12d2f74da1 · outbound

This paper cites A diagram is worth a dozen images.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation A diagram is worth a dozen images

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:17.721711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:34:02.539819Z digest=sha256:c0bb12bec0d6b3a8cbe6a660234771d24d3f137b2f045cd97aee67376d6f5cfa

Observation a52519a8-677c-4275-bdfe-737ef7d2a1d5 · outbound

This paper cites Generating images with multimodal language models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Generating images with multimodal language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:17.490002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:34:02.755733Z digest=sha256:d4f7f8946043486a5f7e76cb8b62d73154b8b75cb1b713f09431de68ce08a148

Observation e870ab70-d91a-408c-ad38-3c090f0f4a8a · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:02.916132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:02.916132Z digest=sha256:0e20f6d3d5092bc83eb6461c5fd989400e7ea563457ba06814f60b5eb4971a7f

Observation 79fde0c5-1a91-4e10-926f-c9d7b402a850 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:03.070205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:03.070205Z digest=sha256:f6ec814e55aa4fca3f8785163ecd00529d186a38270534d30b041fb0eafa0968

Observation 12899f82-3563-44a6-9a78-941a20b1afd3 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:03.206849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:03.206849Z digest=sha256:57e248561087c033a26b1af018ff5e375303d16af469493b0df5e60defbb6a13

Observation 273f10d9-7205-4d87-8284-ca2e3c2ef32a · outbound

This paper cites an unresolved cited work.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:34:17.290173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:34:03.339774Z digest=sha256:80e6e0e44dd5335121e1f91a6b5d02c67f797be1e552398964ca5c5d0266a67b

Observation 9a25378b-887d-4c8f-b888-4c69caf597a7 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:03.476230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:03.476230Z digest=sha256:1ae9d922ebe5bdff58b74794f3951b13f1a91d3688d9e8ba0faa22847eafb87b

Observation e07f2fb3-48c4-46de-91da-d8bf02d63ddb · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Evaluating Object Hallucination in Large Vision-Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:03.678107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:03.678107Z digest=sha256:40d8853301cb10088ec687e3a21929d46789b06bda5f09f38b2b35b9d4ca4387

Observation c4ad44f6-9d62-4bf8-baf6-b5e6f06e999c · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll \' a r, and C.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll \' a r, and C

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:03.799566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:03.799566Z digest=sha256:303a744874f0a57e89e790db5320a85c2b3909eb058bafae65740896eed43c35

Observation 98e74cdd-e023-4e57-82bb-a755b74a2fad · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:03.967311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:03.967311Z digest=sha256:09808fbbbd5ce1f1a35ddf0eb2247d23628213d737fc1c89f98a707c56883aed

Observation f8a241d6-d815-4b89-b308-3c81e11cf694 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Improved Baselines with Visual Instruction Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:04.091274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:04.091274Z digest=sha256:a4b8ae93227d8595d68d6d89168efb7f2ddd8383cef73cdc09b9ae3689466d8a

Observation fc9d2ee5-14e5-4b01-92b2-1e4f746c6f5e · outbound

This paper cites Visual Instruction Tuning.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Visual Instruction Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:04.316929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:04.316929Z digest=sha256:4a8bdd5c96dc2dc78d5bf56241810ea7f7b2d3c2b587368d285d6560f011f313

Observation aa3dcf34-721e-4f12-8389-df32b9d9580c · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:17.057930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:34:04.448654Z digest=sha256:88c5e5a629a98a8ad84be7cc3073f77b6a617fd97e7c6fa7cf58caf61cacd4da

Observation ce5fc821-db12-452d-8655-ac7a47e1b085 · outbound

This paper cites Visual instruction tuning.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Visual instruction tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:04.568020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:04.568020Z digest=sha256:588f8ff7e8b9f0c993a768422bfbcd2df3d8883eb6dd86d567a1078ee9aefe99

Observation 396b95b1-89d1-4a06-9d74-74649dc13acf · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation MMBench: Is Your Multi-modal Model an All-around Player?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:04.747893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:04.747893Z digest=sha256:ed20bb15b9a50e04ae1229f8f96d513abd7a16ba989fb9776372d2deaded08fc

Observation e6ea8826-9138-48fc-b20d-1235829373dc · outbound

This paper cites On the hidden mystery of ocr in large multimodal models, 2024 c.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation On the hidden mystery of ocr in large multimodal models, 2024 c

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:16.931157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:34:04.934626Z digest=sha256:6b0a35e8fdc8c326fee99da2e7df013406e8e424a9fe73a929be699209399b4b

Observation 74ab8d04-505b-4eb6-a6e8-66c7fc7cfcfe · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:05.044523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:05.044523Z digest=sha256:9bd453b2db14bc29b7d2a863167debec777dfcf73fdc3d12742e4a38c38974ca

Observation b37d7701-e2da-4f1e-88e1-4496d035711c · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:05.176742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:05.176742Z digest=sha256:ccb0106eed0ad2b6916f4944aa1ce3cbb607f677ccb1e5c756bb1041f65d51c7

Observation 87ae0daa-a108-4529-bfcc-022c9b0aa4aa · outbound

This paper cites Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:05.315938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:05.315938Z digest=sha256:43e54bcb214100ee745dc2596658b4ace581434bc2142fd6132b74430c25d40d

Observation f1179f3e-2376-42ea-8792-0760c0b39fc7 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Docvqa: A dataset for vqa on document images

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:05.474321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:05.474321Z digest=sha256:458740564e0fbaa82926c191ad2fc85849644f1728d7a99e7e065212d63224cb

Observation f0ddd885-f659-40fb-a918-8f9af08c3c99 · outbound

This paper cites Infographicvqa.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Infographicvqa

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:05.634106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:05.634106Z digest=sha256:371d83a1cb3f3d87f3e1162d7fb4563325b618a414247fa06e7bb3b0a9cb0a76

Observation 2838699e-86a7-441a-b0b1-7da3caea2a41 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:05.794775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:05.794775Z digest=sha256:4bfd3c7a6d535a0565b5eccdd5b8fddf469726ecf2de50ce188101fb13c59fa3

Observation c74be96a-30ec-4c99-aca8-4fa8592c5a13 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Learning transferable visual models from natural language supervision

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:16.730462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:34:05.937934Z digest=sha256:cb38a6230267859f7ef2cd75d08899a6ae052253116139945febf506c076018c

Observation 20385547-d54e-4e6f-ad8d-2dfec6e24874 · outbound

This paper cites an unresolved cited work.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:34:16.439186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:34:06.067878Z digest=sha256:3203baeb6d28f7869b35e94938ff4500bf272fa8ab9f48869533e4ae0bccedd0

Observation 04974964-d317-4823-be2f-d05af2983131 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:06.226877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:06.226877Z digest=sha256:d726fd61be588252c1d723e590673a78134c642a5f74c531fc6b7be9942107b4

Observation 607662ee-0c37-4bf5-a7bb-6f381a55fe69 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation High-resolution image synthesis with latent diffusion models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:06.371337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:06.371337Z digest=sha256:c9043f4698a134d58e51a036746a45938a4b6bf74d7362110e808d4ba831a594

Observation 143f32cf-f12e-4123-b98b-f3ab38e3a9b0 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation High-resolution image synthesis with latent diffusion models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:06.504245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:06.504245Z digest=sha256:eef878ccc416ea0b6cc78c4bb6efc90688a526409ae35ed131b06418dec4b331

Observation 1ee84ecc-3961-4ef5-a0ae-76d81a0a4246 · outbound

This paper cites Multimodal Instruction Tuning with Conditional Mixture of LoRA.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Multimodal Instruction Tuning with Conditional Mixture of LoRA

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:06.624192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:06.624192Z digest=sha256:09db4e7e04956d2be7b112f0e96af9de96157b6ecb1bb339ca970a1bcc562d4c

Observation ec43e49d-4005-475c-9c85-60df80513c47 · outbound

This paper cites Towards vqa models that can read.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Towards vqa models that can read

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:16.178169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:34:06.792015Z digest=sha256:46fd10d9a456c191ebd36de83fe4995ed4b20c9de8a5e422e6664a292016603d

Observation b455ec78-c8a9-44bc-9ca5-25d79365c58f · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:06.956907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:06.956907Z digest=sha256:2d571765cb299a17b8804a4efd1b56b0e50cb99eaeda160be57cd39fbd4e6b0f

Observation b6fb765c-2f9b-4a37-8260-721aab517faa · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:07.113869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:07.113869Z digest=sha256:2d5ec8b82700e17a14e34e07af7f2872227ed053c90c79a17227a34de889a82e

Observation b107e098-ca99-4fa3-9c50-ba1aa94d0951 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Emu: Generative Pretraining in Multimodality

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:07.236721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:07.236721Z digest=sha256:6332d2de98557c2ada764ef528bdc734a38e03a0f34ac6e73b6f667731be5116

Observation 9157b874-4a5e-4133-bc4e-ebce19cfbb08 · outbound

This paper cites Generative multimodal models are in-context learners.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Generative multimodal models are in-context learners

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:15.826498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:34:07.381855Z digest=sha256:a472ec2bfe821c4e04af1cf958e46e4d5a3119d0688fe4911105f00c62202167

Observation b6b0a231-7596-49d1-8b53-45a7d0f0e3ad · outbound

This paper cites CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:07.537541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:07.537541Z digest=sha256:8264b60d1f7b2cd5a0d4a0849b83452590b7953c361e988c66a695b6b37ae981

Observation ff032ef9-9d8c-4eac-9f1c-0536508ec074 · outbound

This paper cites Any-to-any generation via composable diffusion.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Any-to-any generation via composable diffusion

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:15.538723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:34:07.678452Z digest=sha256:23635b41f57ee8f920e927a0a404a0d018d759bd6e7463cb93a9080057182f7c

Observation f3aefd29-efa1-4137-8cf6-000253965e3c · outbound

This paper cites Chameleon: Mixed-modal early-fusion foundation models, 2024.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Chameleon: Mixed-modal early-fusion foundation models, 2024

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:07.871098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:07.871098Z digest=sha256:ff5547212ebde90dd2fffb684b672270e7432cfe262923d657de554598911370

Observation ff2de168-136a-4a9c-a1e1-8385b4d6a2fe · outbound

This paper cites MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:08.000903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:08.000903Z digest=sha256:70b2ff239e964e296ceccd1210ecf63398aa0c335c453e56143c1e51f36bde94

Observation 0129f496-dcae-4c81-bc71-311c6cc4aecf · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:08.124461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:08.124461Z digest=sha256:7be8e767b29ebbe3477ce678de35453b107aa787cecdbe5241cee56dbb18f771

Observation 0b5b9bf6-2cc9-4da9-b2db-d3b84535404f · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:15.221743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:34:08.316611Z digest=sha256:ce26cca688cec0c241fac1599a07f715d1006ec1d92523ec58696c8475f4142c

Observation 969ed406-5d8c-4ea7-9be1-56d92ad587cb · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation LLaMA: Open and Efficient Foundation Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:08.477979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:08.477979Z digest=sha256:1f8a9a5d2882210d199b169e98643bba9f06f5a61ab5b9aa80be9f8be20b15f6

Observation 0ffe2784-43c0-407f-87e9-88ba41cf170c · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:08.648626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:08.648626Z digest=sha256:95517d510d7a4be18dc92643759762515147b8a129f2e0fcb47e83ec98bbb522

Observation 7d658a2d-493f-4c3d-8b4e-c7ef1a5e2a14 · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:08.755303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:08.755303Z digest=sha256:75d15d448c1dc6d8cc2620c86742d91a10d5891e6130e6fef7735c06e5041313

Observation 2219aa29-5f93-4240-afd0-11434893aa1e · outbound

This paper cites OFA: unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation OFA: unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:14.903752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:34:08.885112Z digest=sha256:939832a25558048ba1ea4bc0afdeb573c1dfa7626ea76c0313bb459eda7b7f8f

Observation 7c8c064b-8e4f-45de-9a24-3a57bbc9bddb · outbound

This paper cites Image as a foreign language: BEIT pretraining for vision and vision-language tasks.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Image as a foreign language: BEIT pretraining for vision and vision-language tasks

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.022182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:09.022182Z digest=sha256:a100e5342a23229b1dfec8e01b88a652a12d893bd281081a20f01b4939e1eeed

Observation 1d99262d-5b9d-4c38-be12-0575e788c3eb · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Emu3: Next-Token Prediction is All You Need

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.192768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:09.192768Z digest=sha256:7a284c8e420fac87b961ec14cc47ab291a41e6a21d4888a5a5e3618c98819d08

Observation 974ce6fa-37c5-4c72-883f-1f57d968f342 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.374451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:09.374451Z digest=sha256:bd70e2353ddc70194b7559e58f477978a9ca84228127e6ec47da89f9ce68e045

Observation d18c754e-cc81-41f7-95a0-bf11ed74dfb9 · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.517877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:09.517877Z digest=sha256:1746b346321c40b18cfe396ea9ba1885d6b4e3b7468172dddf1bae0024ab247a

Observation aa67e46c-d63e-481c-8fef-55eceba2b082 · outbound

This paper cites Grok 1.5v: The next generation of ai.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Grok 1.5v: The next generation of ai

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:14.727112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:34:09.703995Z digest=sha256:8a4790ea6b7ed352702568ba68ac0314863168f8cb979573ac097121a9b16eed

Observation ff7c1f67-5593-4f19-ae75-b98806a9f200 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.844460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:09.844460Z digest=sha256:073f2ab32ae82fa86e6d4e23122194eba01003e30d4165b5e3eddc8599d7d3c7

Observation d30a2810-ca8f-4506-841d-cab57a499d09 · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.966497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:09.966497Z digest=sha256:6bf38b44d135f996b6c4b4b5bb9bc0ee8b8f60b29cc005f9b417957b9cd75c6b

Observation 877fdd3f-effc-4341-8f7d-aee0f9d1ff3a · outbound

This paper cites Multiinstruct: Improving multi-modal zero-shot learning via instruction tuning.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Multiinstruct: Improving multi-modal zero-shot learning via instruction tuning

Reference 73

Resolution
verified exact
doi, observed 2026-08-07T04:34:12.568692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:34:10.109785Z digest=sha256:3d331e28e12a27e29b9978e700011a51b906b85afc6737196432d7bec26634c5

Observation 60240428-c3d7-42de-b785-8280a413f474 · outbound

This paper cites Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:10.263657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:10.263657Z digest=sha256:2ee81a76b7e1d0594783b1e036544f53d38379d722323d32f43a9f5ac2359831

Observation 1cd2c1f0-cfcb-4753-a007-e055acbfb868 · outbound

This paper cites Modality-specialized synergizers for interleaved vision-language generalists.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Modality-specialized synergizers for interleaved vision-language generalists

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:14.442097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:34:10.391353Z digest=sha256:0ff6d45a52386f3a25522a9b124357567b913e5115caed2bdd981d1216688e99

Observation 8f22ff07-7612-4a2a-b947-666c44328754 · outbound

This paper cites Retrieval-augmented multimodal language modeling.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Retrieval-augmented multimodal language modeling

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:14.215441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:34:10.582907Z digest=sha256:07b57b692617b11c22f030449ec7ad7240a326a270ff1fa09fa714cbd5d113b9

Observation e7c6d5cf-9f27-42b2-ad58-643a88ed7b46 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:10.749553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:10.749553Z digest=sha256:2dca81ad596afbbd199ecc55017d31935fd61ea4bbb8b5a1f0a1fd2208a52c19

Observation eb2dfaed-8279-4d42-9cac-c2fa98baea49 · outbound

This paper cites LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:10.895265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:10.895265Z digest=sha256:d92552d6b1d65de01fb100777133192b0fb4ee67a9b8a84189883555bc12210a

Observation e9b4edd9-af5c-4184-8228-db135bf16a19 · outbound

This paper cites Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:11.061508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:11.061508Z digest=sha256:d84471d59955a41562bc4cfa4217e3f1e1ad8c7995927cebe6829435d96ce495

Observation 825db274-4934-4ccb-9e22-98aa3e98a425 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:11.250107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:11.250107Z digest=sha256:4a49072c641ce9e9f24c19c16541b29584ed6e1ec2167eb30685adaf4268b0a9

Observation 12e13317-6f53-4ab6-ad1b-3267e3526906 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:11.405831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:11.405831Z digest=sha256:a575fd734fc259a5c6c998b730b4389153590801db21f499f1dcb1a5cebff99f

Observation 1f35042e-3a44-45ce-86a2-9093cc06b9a2 · outbound

This paper cites Sigmoid loss for language image pre-training.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Sigmoid loss for language image pre-training

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:11.533516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:11.533516Z digest=sha256:5d814a796a3682274e7e09c42f8bae7ed2c3e31672fbe875c00ca273240a92f0

Observation c35cc7c8-52e8-4e1f-897b-b484e3188537 · outbound

This paper cites AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:11.720293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:11.720293Z digest=sha256:72c23ebd029a8234d26e0615cf45c66a4e52e4425372e0b3c6b05a9bedc69645

Observation 1137c489-997b-411f-82db-76314dc056c9 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:11.878440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:11.878440Z digest=sha256:d8c677c623dc1faae0dda4979d34857347dd78d5b03786813f002fc30b32cca3

Observation df24cd45-8a6c-4865-9926-9dcdf3affe65 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:12.006826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:12.006826Z digest=sha256:66b6aa5291302c1dedd944b46961b9190aad7f15d1f82d12c8574906977637d1

Observation 9c091c59-76f3-4992-bda8-50607a9d87c7 · outbound

This paper cites VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:12.154340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:12.154340Z digest=sha256:c07a369ce1ebdc2b86134819c9a75e2959483ac8c6c3106c80e266ba4f9512e7

Pith citing papers

Observation 00f45ee6-cf2c-4edd-b665-fcf19c93116a · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation

Reference 132

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.884365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:07d6bcb4efd559b1666f2ab380273ff57c41f0bea70a21617aa6fb3e565ee63d

Observation 98aa2f0d-0924-40d1-a253-9fb7348f8004 · inbound

Transferability Between Understanding and Generation in Unified Multimodal Models cites this paper.

Transferability Between Understanding and Generation in Unified Multimodal Models Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation

Reference 97

Resolution
unresolved
no resolver link, observed 2026-07-11T19:17:19.634242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:17:19.634242Z digest=sha256:24a33e5a7c8c3f622d86bb8770cd04b508508a3191de3e350ab9e0213f990d84