REVIEW 5 cited by
T-Rex: Counting by Visual Prompting
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce T-Rex, an interactive object counting model designed to first detect and then count any objects. We formulate object counting as an open-set object detection task with the integration of visual prompts. Users can specify the objects of interest by marking points or boxes on a reference image, and T-Rex then detects all objects with a similar pattern. Guided by the visual feedback from T-Rex, users can also interactively refine the counting results by prompting on missing or falsely-detected objects. T-Rex has achieved state-of-the-art performance on several class-agnostic counting benchmarks. To further exploit its potential, we established a new counting benchmark encompassing diverse scenarios and challenges. Both quantitative and qualitative results show that T-Rex possesses exceptional zero-shot counting capabilities. We also present various practical application scenarios for T-Rex, illustrating its potential in the realm of visual prompting.
Forward citations
Cited by 5 Pith papers
-
DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts
Global prompt integration, visual-textual relation distillation and selective fusion make visual prompts discriminative enough for DETR-ViP to beat prior visual-prompt detectors by several mAP points.
-
DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models
DINO-R1 trains visual-prompt detectors with group-relative query rewards and KL regularization, improving zero-shot and fine-tuned detection over supervised fine-tuning.
-
Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning
Training a multimodal LLM on GPT-4o-generated chain-of-thought referring traces, then optimizing with GRPO, improves referring accuracy and abstention on HumanRef.
-
Expanding Zero-Shot Object Counting with Rich Prompts
RichCount improves zero-shot object counting by enriching text prompts with MLLM-generated descriptions and aligning them to CLIP visual features, achieving state-of-the-art mean absolute error on three counting benchmarks.
-
Cotton-SF YOLO: Learning Structural and Frequency Cues for Early Cotton Square Detection in Complex Field Environments
Adding dynamic-snake-convolution and FFT-modulation modules to YOLO26m raises cotton-square detection mAP50 from 0.810 to 0.820, mAP50:95 from 0.478 to 0.494, and recall from 0.771 to 0.794 on a new field dataset.
Discussion (0). Continue with ORCID to comment.