Pith. sign in

Latent-space scalability for multi-task collaborative intelligence

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We investigate latent-space scalability for multi-task collaborative intelligence, where one of the tasks is object detection and the other is input reconstruction. In our proposed approach, part of the latent space can be selectively decoded to support object detection while the remainder can be decoded when input reconstruction is needed. Such an approach allows reduced computational resources when only object detection is required, and this can be achieved without reconstructing input pixels. By varying the scaling factors of various terms in the training loss function, the system can be trained to achieve various trade-offs between object detection accuracy and input reconstruction quality. Experiments are conducted to demonstrate the adjustable system performance on the two tasks compared to the relevant benchmarks.

citation-role summary

background 1

citation-polarity summary

fields

cs.CV 1

years

2026 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

background 1

representative citing papers

Visual Token Codec: Unleashing Spatial Redundancy for ViT Feature Coding

cs.CV · 2026-08-09 · conditional · novelty 6.0

The Visual Token Codec compresses ViT intermediate features by entropy-coding patch tokens on their native grid instead of a flattened sequence, reducing bitrate by 15.7x to 37.4x at 90% of uncompressed performance over a VTM-based baseline.

citing papers explorer

Showing 1 of 1 citing paper.

  • Visual Token Codec: Unleashing Spatial Redundancy for ViT Feature Coding cs.CV · 2026-08-09 · conditional · none · ref 53 · internal anchor

    The Visual Token Codec compresses ViT intermediate features by entropy-coding patch tokens on their native grid instead of a flattened sequence, reducing bitrate by 15.7x to 37.4x at 90% of uncompressed performance over a VTM-based baseline.