Pith. sign in

Paper Citation Record · LEDGER

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA

As of 19 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2505.06356.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.06356 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:47:49.169054Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 663b3fed-252e-4bcd-9f2b-3b31b2a02610 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.579507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:47:49.049971Z digest=sha256:35715f2af2c99dd55fe39cfcf16fe66bcc080f48184a755dc5337c2266b6a50f

Observation cbaf101d-231b-4de9-9e3a-656974b9fcbb · outbound

This paper cites Y our vision-language model itself is a strong filter: Toward s high-quality instruction tuning with data selection, 2024.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Y our vision-language model itself is a strong filter: Toward s high-quality instruction tuning with data selection, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.566146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:47:49.054656Z digest=sha256:b518afeb955bce2c2c5a82c0dee3d1ff73cdbcc39dd898759ef3c3a614562d83

Observation c8739de5-0843-4a1d-9212-26007803a608 · outbound

This paper cites Comm: A coherent inter- leaved image-text dataset for multimodal understanding an d generation, 2024.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Comm: A coherent inter- leaved image-text dataset for multimodal understanding an d generation, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.553590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:47:49.058646Z digest=sha256:558d347182a17f07cc375e5b40e28c7c00775a791d6842fac9ca2a29fdeb52a7

Observation 571a04cb-9f43-4929-ab0f-3c26ca496045 · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T22:47:49.063146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:47:49.063146Z digest=sha256:ebdbcccd49bd3a3e9cd4cc33d106a480210fc00f11bf71691aaa7e74c5c6f092

Observation beb628dc-a79c-441d-b713-aeae39b9511f · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T22:47:49.067724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:47:49.067724Z digest=sha256:44eeb716da29ab3d5ee743687d631b0b55e029f03c79ac3e5b228967ce04a6cc

Observation 4b1fa9a7-9aa9-4813-9bd3-f003bb63ac61 · outbound

This paper cites Command R.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Command R

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.540544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:47:49.072264Z digest=sha256:7f7706ef44911595c93fcce64d3e44df355250374d88f9cb060a3d90f864ec70

Observation 4b29b4f1-c834-4464-a24f-ca449b161f2c · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T22:47:49.076972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:47:49.076972Z digest=sha256:c6da01f67583fe7a3cd91030d7fef05c18dc3d0c4898cb076bd9e86a6037f1c8

Observation 11a0ef41-deff-45d3-befd-cdea1cce37d2 · outbound

This paper cites Detoxify.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Detoxify

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.527155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:47:49.081229Z digest=sha256:8ae4e04aabc0ef95008f6ce59554cbf15da6197c06157d236b17d007222f2bf1

Observation b1b27d42-1677-40a9-80d5-d26b7f492402 · outbound

This paper cites LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T22:47:49.084993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:47:49.084993Z digest=sha256:3d4262c16a90597cab9d9461abb4f39a73ca2e6f1dbdfb51c93a995eae9dd811

Observation 28400e92-8541-4288-8602-28fc2a78a483 · outbound

This paper cites Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language an d vision-language models, 2024.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language an d vision-language models, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.513777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:47:49.089163Z digest=sha256:b492ec047fdb124f0541c9d66ece6057c58928b488c46fe1df433045909593e1

Observation 72617fd2-c0ba-45ae-8ca4-8d833f511516 · outbound

This paper cites Vhelm: A holistic evaluation of vision language mod- els, 2024.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Vhelm: A holistic evaluation of vision language mod- els, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.500892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:47:49.092974Z digest=sha256:5fe64ce80887fbff042e5e7812575c9d8f3a3eefb87f745691e018a4336b908b

Observation 0016f5da-6da0-4bd6-832d-0f41ca6cdf4f · outbound

This paper cites Elite: Enhanced language-image toxicity evaluation for safety, 2025.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Elite: Enhanced language-image toxicity evaluation for safety, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.487453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:47:49.096768Z digest=sha256:31a023e34bd7e1c00b4c5786fb2f58587fe11cabc236830591369c5d4c661462

Observation 7c52c311-f8aa-44e2-982b-3c03ae64f2f2 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning, 2023.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Improved Baselines with Visual Instruction Tuning, 2023

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.472673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:47:49.100663Z digest=sha256:e4e34516417d7261181688a383a770b0099687d1fa1f8a01449b8fd04fb568ad

Observation 0f0d46d8-289f-45b3-80c1-1269dd77d017 · outbound

This paper cites Visual Instruction Tuning, 2023.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Visual Instruction Tuning, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.458294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:47:49.104582Z digest=sha256:c14c7336e81165efb9e7f3e184de3f4120ab4715981f9ab626b6ea81c7ececfe

Observation 13475d53-5d5d-486c-b399-112d1f2cb2f0 · outbound

This paper cites Mm-safetybench: A benchmark for safety eval- uation of multimodal large language models, 2024.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Mm-safetybench: A benchmark for safety eval- uation of multimodal large language models, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.445746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:47:49.108332Z digest=sha256:378ce8123c5a327a6165612a6d79edda7869bc916883e0573077f128fc42225e

Observation 2a99a062-e29d-44f4-b738-b0a4e56e1c93 · outbound

This paper cites Towards interpreting visual infor - mation processing in vision-language models, 2024.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Towards interpreting visual infor - mation processing in vision-language models, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.431701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:47:49.112379Z digest=sha256:9c45fcf7df89c4f5b91bdcbe687849036acd643e118efc7f3228e623c006f790

Observation 92697fc4-b12a-4c82-99dd-ea89b95b9e8a · outbound

This paper cites Kosmos-G: Generating Images in Context with Multimodal Large Language Models.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T22:47:49.116160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:47:49.116160Z digest=sha256:507e52d387708ae8a9a52b7a36cc98fc8cbaa47589025bb27113ca46f61edc0d

Observation 1647ad66-906c-4a99-940b-0121ea53c54e · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T22:47:49.120423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:47:49.120423Z digest=sha256:eebb5c8b2d6f45148d4e09edeb36e324cdf4f2f7f2de26e075419ccf2d0a794c

Observation 088b5361-2522-4ad9-b26a-8b90d41c74a5 · outbound

This paper cites Learn- ing Transferable Visual Models From Natural Language Su- pervision.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Learn- ing Transferable Visual Models From Natural Language Su- pervision

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.418176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:47:49.124792Z digest=sha256:a38d6a544fdd72ac075e8fddd4e8c1b05fea7cbd189aaecac30af043c18eeeac

Observation 6543f544-6704-401b-9c6b-d4d5116611c4 · outbound

This paper cites Training-free mitigation of language reasoning degradation after multimodal instruction tuning, 2024.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Training-free mitigation of language reasoning degradation after multimodal instruction tuning, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.404586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:47:49.130340Z digest=sha256:ae6677caff032c89944f36c12610e24b72dc88a1e96a26962d73eb78210263f6

Observation 852ddd3e-c245-4c8d-80de-3d3a7fbc7846 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models, 2022.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Laion-5b: An open large-scale dataset for training next generation image-text models, 2022

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.391510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:47:49.134378Z digest=sha256:1dc50874bc34291a0fce10bfa64162be25a604a5de93992ebe708ad5196cb4f9

Observation 66fd9b2b-a4bd-4769-a8c1-b8851402f569 · outbound

This paper cites From pixels to prose: A large dataset of dense image cap- tions, 2024.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA From pixels to prose: A large dataset of dense image cap- tions, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.378068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:47:49.138090Z digest=sha256:e5928dbc935483ceb443074566698a344dd538bd6da28a9ae55ea26892335e56

Observation 0e45e7a4-cd81-4d42-b2c9-df0cac1f1b5f · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding, 2021.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA RoFormer: Enhanced Transformer with Rotary Position Embedding, 2021

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T22:47:49.141758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:47:49.141758Z digest=sha256:e3ba989a63a54a4b9fe67c34769898af82715f036a1d120996201dbaecf0adda

Observation e120afa9-4647-45ad-afdd-aae5e0c7bdef · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T22:47:49.145359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:47:49.145359Z digest=sha256:2996f2e8faf78e17739feae10838c0ae86069b13a4b90f4cd3673fc9c8fb05b6

Observation 1aeacf78-2ebc-4b92-840c-ed201d70f13a · outbound

This paper cites Florence-2: Advancing a unified representation for a variet y of vision tasks.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Florence-2: Advancing a unified representation for a variet y of vision tasks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T22:47:49.148914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:47:49.148914Z digest=sha256:21e8fae7fa324bd61c3d9342a810f41b7f47f983f5715a5cee262a9ae57b2635

Observation 3da6cc34-7eb3-48ab-9aca-e841e01b8453 · outbound

This paper cites Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T22:47:49.152993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:47:49.152993Z digest=sha256:bcbe12f1241c7bcee64828a9c8ad685d7507cef49e6bd8dc8fe42569f29f8f43

Observation 4f8a7d9c-5f57-4835-9f93-a60c693cfe9e · outbound

This paper cites Sigmoid Loss for Language Image Pre- Training.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Sigmoid Loss for Language Image Pre- Training

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.348897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:47:49.157150Z digest=sha256:3d5b5cb4eabcd4d59147ea8856098095690d06206afa8a10cd8f02501e0cd403

Observation 23d66c40-6f80-4e29-a430-85d07664c789 · outbound

This paper cites Spa-vl: A comprehensive safety preference alignment dataset for vi - sion language model, 2025.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Spa-vl: A comprehensive safety preference alignment dataset for vi - sion language model, 2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.336125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:47:49.161173Z digest=sha256:a390f3d706a489fc415682ecd95bb879ec657752c70d4d347ec9c3f05f3dc88e

Observation aeeabb02-1824-46e6-aa80-75e3ffa4eb70 · outbound

This paper cites Zero-shot defense against toxic images via inherent multimodal alignment in lvlms, 2025.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Zero-shot defense against toxic images via inherent multimodal alignment in lvlms, 2025

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.322869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:47:49.165134Z digest=sha256:a7054fd86360e7151fab35708b27173a73938b53d65be38ec6c8dea5afeb4cff

Observation 53243062-5057-4e78-81f9-1e3cb69f06cd · outbound

This paper cites Un- derstanding and rectifying safety perception distortion i n vlms, 2025.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Un- derstanding and rectifying safety perception distortion i n vlms, 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.309450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T22:47:49.169054Z digest=sha256:1a5d56e41d1e1ae2c4e14fbd2c83c7708b85005c76ae371a15332ac4e112628b

Pith citing papers

No inbound Pith citation observations are available.