Pith. sign in

hub Canonical reference

M3it: A large-scale dataset towards multi-modal multilingual instruction tun- ing

Canonical reference. 83% of citing Pith papers cite this work as background.

11 Pith papers citing it
Background 83% of classified citations

hub tools

citation-role summary

background 4 dataset 1 method 1

citation-polarity summary

representative citing papers

Vision-Language Foundation Models as Effective Robot Imitators

cs.RO · 2023-11-02 · conditional · novelty 6.0

RoboFlamingo adapts open-source vision-language models for robot manipulation tasks via single-step comprehension plus an explicit policy head, outperforming prior methods on benchmarks with only light fine-tuning.

Large Language Models are not Fair Evaluators

cs.CL · 2023-05-29 · conditional · novelty 6.0

LLMs show strong position bias when scoring model outputs, allowing easy manipulation of rankings, but calibration with multiple evidence, position balancing, and selective human input reduces this bias to better match human judgments.

Otter: A Multi-Modal Model with In-Context Instruction Tuning

cs.CV · 2023-05-05 · unverdicted · novelty 6.0

Otter is a multi-modal model instruction-tuned on the MIMIC-IT dataset of over 3 million in-context instruction-response pairs to improve convergence and generalization on tasks with multiple images and videos.

A Survey on Multimodal Large Language Models

cs.CV · 2023-06-23 · accept · novelty 3.0

This survey organizes the architectures, training strategies, data, evaluation methods, extensions, and challenges of Multimodal Large Language Models.

citing papers explorer

Showing 11 of 11 citing papers.