Pith. sign in

hub Mixed citations

Qwen2.5-vl technical report

Mixed citation behavior. Most common role is background (33%).

12 Pith papers citing it
Background 33% of classified citations

hub tools

citation-role summary

background 2 dataset 2 method 2

citation-polarity summary

years

2026 11 2025 1

verdicts

UNVERDICTED 12

representative citing papers

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm

cs.CV · 2026-05-12 · unverdicted · novelty 7.0 · 2 refs

Proposes V2V-Zero, a training-free framework replacing text conditioning with VLM final-layer hidden states from visual pages, achieving 0.85 on GenEval and 32.7/100 on new Simple-V2V Bench across models including video extension.

HunyuanVideo 1.5 Technical Report

cs.CV · 2025-11-24 · unverdicted · novelty 6.0

HunyuanVideo 1.5 delivers state-of-the-art open-source text-to-video and image-to-video generation with an 8.3B parameter DiT model featuring SSTA attention, glyph-aware encoding, and progressive training.

Do Vision Language Models Need to Process Image Tokens?

cs.CV · 2026-04-10 · unverdicted · novelty 5.0

Visual representations in VLMs converge quickly to stable low-complexity forms while text continues evolving, with task-dependent needs for sustained image token access.

citing papers explorer

Showing 12 of 12 citing papers.