Pith. sign in

Paper Citation Record · LEDGER

Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2408.15556.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.15556 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:40:40.270764Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T11:46:55.941528Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 498932a9-8030-4f1a-b15d-e3f277ec51ff · inbound

Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs cites this paper.

Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:35:13.196401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T05:35:13.118221Z digest=sha256:0a881d122de45016cf22579f115eac260e511e50c242cbdb9b7eb396e3b19ef3

Observation fd65673c-bd1a-4940-8b85-dd4bc318fbda · inbound

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images cites this paper.

A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:40:40.270764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:40:40.270764Z digest=sha256:fbe07d0db2e5694733d9509c0dc48b076d1a88bc8cc90599f055866c80cd01ed

Observation 3aeafd0e-95ae-4f16-88e7-3beea45346ff · inbound

SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring cites this paper.

SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:16.703802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T16:48:20.100672Z digest=sha256:7310048b6b2ad892b27537a74fd429df917544b5cd7c27b3c91bf9b0cfea46db

Observation cfccd4ef-e3ed-4276-92c4-9958e26dc9f4 · inbound

SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring cites this paper.

SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:19:50.048642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T06:15:55.968506Z digest=sha256:0a6e5f682bd6ab6fad0417c6e3d2fbf0b075c4b2a8f447f9937ec01c607ce359

Observation a1e7dca6-0b09-4fca-bcf0-1a96bd5b4578 · inbound

MCMit: Mid-Circuit Measurement Error Mitigation cites this paper.

MCMit: Mid-Circuit Measurement Error Mitigation Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:45:35.071706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T08:36:37.641628Z digest=sha256:5f7d6129df712ac2690b88251f345a44a2ac0332e4c24db410480a24bed4efac

Observation abc9bb63-8f73-44f8-9c58-219f56106c5c · inbound

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth cites this paper.

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:53:13.083496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T10:52:32.238241Z digest=sha256:bbd04790a6c851bb017b074181716942cfa3c86a079369d1a38ba788b22fd684

Observation ae552715-b7ea-4c3d-a7e2-005de79615fa · inbound

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth cites this paper.

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T16:34:34.452430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T16:34:34.452430Z digest=sha256:e9e7880c7c0360d574f8c2f449c192f8fb41623c5baed27ab225bccf0305929b

Observation 94de064a-f031-4d7e-8ae6-45190a185b19 · inbound

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison cites this paper.

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:39:53.717075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T08:36:27.676888Z digest=sha256:c9317274973e679fd94057efaa2d26c9ee47e54a87e50d1ea77f6da056594979

Observation f5651ffc-427a-40e7-877f-47917c8181a4 · inbound

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison cites this paper.

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:04:58.341749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T17:57:47.409741Z digest=sha256:91ff953c448202f118a135f829670108e0cff27dfcc601571e4e36bbfb320883

Observation 2b012226-e0ae-4d93-9408-3da5044fcadc · inbound

UltraVR: A Diagnostic Ultra-Resolution Image-VQA Benchmark for Evidence-Grounded Reasoning cites this paper.

UltraVR: A Diagnostic Ultra-Resolution Image-VQA Benchmark for Evidence-Grounded Reasoning Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:46:55.943105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T02:58:25.590904Z digest=sha256:7551f4fda3fffdd3bfec39b8009e6a9529e75508f8f33a0ae8f4e388155a0b85

Observation fa4d5bd5-276d-4984-985d-7ee87fed7d0a · inbound

AdaTurn: Budget-Aware Test-Time Scaling for Active Visual Perception Agents cites this paper.

AdaTurn: Budget-Aware Test-Time Scaling for Active Visual Perception Agents Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T01:50:53.323606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:50:53.323606Z digest=sha256:f1f11af794f2108841a2fa5f206586c82e0d706715266d88542dd14132202942

Observation f14c760b-b865-4160-9267-8a621f6c0ccf · inbound

RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models cites this paper.

RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-31T14:46:30.136196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T14:46:30.136196Z digest=sha256:631da06580249726a8bae3a07ed4990fd7cc51185ff015e39423b13b348abbeb

Observation aa723826-0f6f-48dd-8d98-7c120e7a0deb · inbound

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation cites this paper.

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-07-31T02:56:58.284636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T02:56:58.284636Z digest=sha256:6e12e2bbaa207edcb97616fbc9ecd319a79b0b3e53604f9eaa78c5d73ba74834