Pith. sign in

Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd, 2024

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.CV 1

years

2024 1

verdicts

CONDITIONAL 1

representative citing papers

citing papers explorer

Showing 1 of 1 citing paper.

  • Sharingan: Extract User Action Sequence from Desktop Recordings cs.CV · 2024-11-13 · conditional · none · ref 5

    Directly feeding sampled desktop-recording frames to a vision-language model extracts click/select/scroll/drag/type action sequences more reliably than explicitly computing frame differences first.