Pith. sign in

Paper Citation Record · LEDGER

Detecting Text Manipulation in Images using Vision Language Models

As of 7 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2509.10278.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.10278 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T18:04:20.574759Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-03T20:02:56.057102Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:08:55.138666Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 70313364-3ac7-4635-bfda-c39daa62fee4 · outbound

This paper cites Qwen2.5-VL Technical Report.

Detecting Text Manipulation in Images using Vision Language Models Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:18.054749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:18.054749Z digest=sha256:da374ca38ffc00d28eef5a15da5cedbbb362e4a32fee8a13c5a0383f6e2ab8c7

Observation 9c311dd9-9dc4-4572-9b57-9db1271ac223 · outbound

This paper cites Textdiffuser- 2: Unleashing the power of language models for text rendering.

Detecting Text Manipulation in Images using Vision Language Models Textdiffuser- 2: Unleashing the power of language models for text rendering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:18.254757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:18.254757Z digest=sha256:c98d3c40ec3fb59f97274918e65c5e5851854afcc801ae16d7b0715856ca4d83

Observation 2053abec-4f97-4f26-b31d-20405ddcceb2 · outbound

This paper cites Image manipulation detection by multi-view multi-scale supervision.

Detecting Text Manipulation in Images using Vision Language Models Image manipulation detection by multi-view multi-scale supervision

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:18.425371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:18.425371Z digest=sha256:1d94c0d52ee56a8542026d997d07ff673f20de7115725f3eee1b948fbdca8d63

Observation af2f30ca-4554-4e2a-9bc9-d1c5b424d0b3 · outbound

This paper cites Chatbot arena: An open platform for evaluating llms by human preference.

Detecting Text Manipulation in Images using Vision Language Models Chatbot arena: An open platform for evaluating llms by human preference

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:18.518832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:18.518832Z digest=sha256:86afbecd6dd6941de02abdface6e91b9415f1432947ba84ac8670b8e35ed842f

Observation 67b6ed52-900b-46f3-8cb4-77cb1ff70fbf · outbound

This paper cites On the detection of digital face manipulation.

Detecting Text Manipulation in Images using Vision Language Models On the detection of digital face manipulation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:18.632298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:18.632298Z digest=sha256:7d523d0b8c9d406d19837879f6f05ac614de7743656bff0921df735e17207192

Observation f10857ed-5151-46ae-9b21-d3ca7fa3f3a6 · outbound

This paper cites Mvss-net: Multi-view multi- scale supervised networks for image manipulation detection.IEEE Transactions on Pattern Anal- ysis and Machine Intelligence, 45(3):3539–3553, 2022.

Detecting Text Manipulation in Images using Vision Language Models Mvss-net: Multi-view multi- scale supervised networks for image manipulation detection.IEEE Transactions on Pattern Anal- ysis and Machine Intelligence, 45(3):3539–3553, 2022

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:18.712784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:18.712784Z digest=sha256:e5b2b7a3d7c57f6e2481fc27238a2008fc364994cd839814fa70ad1dd6d57816

Observation 9d389dfa-c654-422d-acde-acf72991fc14 · outbound

This paper cites Casia image tampering detection evaluation database.

Detecting Text Manipulation in Images using Vision Language Models Casia image tampering detection evaluation database

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:19.004846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:19.004846Z digest=sha256:23970f75e42f119c9662047f8d12704602eb730d11a1d16e3b675cbf1033f56d

Observation ee526ebb-f2ea-4c8b-9b16-f86c92c165db · outbound

This paper cites AMMeBa: A Large-Scale Survey and Dataset of Media-Based Misinformation In-The-Wild.

Detecting Text Manipulation in Images using Vision Language Models AMMeBa: A Large-Scale Survey and Dataset of Media-Based Misinformation In-The-Wild

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:19.084829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:19.084829Z digest=sha256:7d7445e81a1324746bca016fbf76d28abf03925a16d2853926cbc44a872e2e30

Observation 3be553d2-7ed7-42da-8b78-37e74bb50334 · outbound

This paper cites The Llama 3 Herd of Models.

Detecting Text Manipulation in Images using Vision Language Models The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:19.213306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:19.213306Z digest=sha256:826201babf22d8b753e06d2fb857e1d3ceb47d259db94ed07676b97f4f807341

Observation 364d1e11-953a-4e0f-8f92-19b467d019fd · outbound

This paper cites Tru- for: Leveraging all-round clues for trustworthy image forgery detection and localization.

Detecting Text Manipulation in Images using Vision Language Models Tru- for: Leveraging all-round clues for trustworthy image forgery detection and localization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:19.444751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:19.444751Z digest=sha256:024ad40954cc95d584349cacd07d1a17b8d3d1579039bbabddbc81553549942c

Observation de5a0f2b-6b45-472e-8df7-e002ff5e1642 · outbound

This paper cites Hier- archical fine-grained image forgery detection and localization.

Detecting Text Manipulation in Images using Vision Language Models Hier- archical fine-grained image forgery detection and localization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:19.577394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:19.577394Z digest=sha256:aa75c066b3b36f66a341bc04b93ecbabbb27b7753b90c5871f1284a8bac5be5c

Observation 8bc15881-3022-4427-8dcc-7fea0464ccfd · outbound

This paper cites SIDA: Social Media Image Deepfake Detection, Localization and Explanation with Large Multimodal Model.

Detecting Text Manipulation in Images using Vision Language Models SIDA: Social Media Image Deepfake Detection, Localization and Explanation with Large Multimodal Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:19.694757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:19.694757Z digest=sha256:91eb0de8cc66f7654c0029ef8532cc83667d21114c8a6db45c5bca873b71d125

Observation 319047f5-1dad-4a24-b3e9-5dafa75af0af · outbound

This paper cites The point where reality meets fan- tasy: Mixed adversarial generators for image splice detection.Advances in neural information processing systems, 32, 2019.

Detecting Text Manipulation in Images using Vision Language Models The point where reality meets fan- tasy: Mixed adversarial generators for image splice detection.Advances in neural information processing systems, 32, 2019

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:19.870008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:19.870008Z digest=sha256:bd1a2d04195bb6897e0426c80732209cc5a3c03722006d75d6b05d925f7877cd

Observation 06f360f2-0ac5-4bdd-8eb3-c4d96ac76b3d · outbound

This paper cites Exploring chatgpt for face presentation attack detection in zero and few-shot in-context learning.

Detecting Text Manipulation in Images using Vision Language Models Exploring chatgpt for face presentation attack detection in zero and few-shot in-context learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:20.064836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:20.064836Z digest=sha256:ffe5e9c333ed2784319db8767d6ab20dc39c285515ccda12adbe6563cbe5ae3b

Observation fb3976a8-f6d0-4f6d-a82f-60e8bfa54103 · outbound

This paper cites FantasyID: A dataset for detecting digital manipulations of ID-documents.

Detecting Text Manipulation in Images using Vision Language Models FantasyID: A dataset for detecting digital manipulations of ID-documents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:20.314744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:20.314744Z digest=sha256:d92d4d6fdb1cb08eec295d6ed91b561e5a7f3314ad95fa764629474c3942933a

Observation 99a6db89-cbf4-4995-9191-9f09dc58bc09 · outbound

This paper cites Cat-net: Compression artifact tracing network for detection and localization of image splicing.

Detecting Text Manipulation in Images using Vision Language Models Cat-net: Compression artifact tracing network for detection and localization of image splicing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:20.437130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:20.437130Z digest=sha256:359296db485d45fdd2436116a8cf042dfb2347eccb4e797bcbc5f10260cb1491

Observation 81f44cdf-b203-4b89-aef2-6b9571f8938c · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Detecting Text Manipulation in Images using Vision Language Models Lisa: Reasoning segmentation via large language model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:20.574759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:20.574759Z digest=sha256:e3d19b407a537a130db16d350a54a3a6cacca0bbdcb6fa8e7270279d7a8961c8

Pith citing papers

Observation 26fe4bc7-e59a-4e4e-acee-314a42cff1f3 · inbound

From Forgeries to Foundation Models: A Systematic Survey of Identity Document Attack and Detection cites this paper.

From Forgeries to Foundation Models: A Systematic Survey of Identity Document Attack and Detection Detecting Text Manipulation in Images using Vision Language Models

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:08:55.141276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T20:02:56.057102Z digest=sha256:0dbfbfb0c32c7a570aefe8faf7f0bfac9346d7b4425cc39e9086161ce14eba47