Pith. sign in

hub Canonical reference

Masked image pretraining on language assisted representation

Canonical reference. 80% of citing Pith papers cite this work as background.

57 Pith papers citing it
156 external citations · external index
Background 80% of classified citations

hub tools

citation-role summary

background 12 baseline 2 method 1

citation-polarity summary

representative citing papers

Brain-IT-VQA: From Brain Signals to Answers

cs.CV · 2026-05-28 · unverdicted · novelty 7.0

Brain-IT-VQA decodes visual question answers from fMRI using a transformer to extract language tokens and introduces the NSD-VQA benchmark with 20 controlled questions per image across 20 categories.

Evaluating the Search Agent in a Parallel World

cs.AI · 2026-03-05 · unverdicted · novelty 7.0

Mind-ParaWorld creates parallel worlds with atomic facts to evaluate search agents on future scenarios, showing they synthesize evidence well but struggle with collection, coverage, sufficiency judgment, and stopping decisions.

Symbolic recovery of PDEs from measurement data

cs.LG · 2026-02-17 · unverdicted · novelty 7.0

Symbolic rational-function networks recover an admissible PDE from noiseless complete measurements and select the regularization-minimizing parameterization within the architecture.

Seed-Guided Semi-Supervised Clustering by A-Contrario Anomaly Detection

cs.LG · 2026-06-17 · unverdicted · novelty 6.0

Introduces the Perception algorithm for seed-guided semi-supervised clustering via a-contrario anomaly detection, defining clusters as anomaly-free subsets and achieving competitive performance with 10-30 seeds per cluster on benchmarks.

Quaternion Self-Attention with Shared Scores

cs.LG · 2026-05-24 · unverdicted · novelty 6.0

Shared-score quaternion self-attention reduces score multiplications by 75% and softmax operations from four to one while proving equivalence to component-wise attention under quaternion linear projections.

Rubato: Transcribing Piano Music with Timestamps

cs.SD · 2026-05-22 · unverdicted · novelty 6.0

Rubato model with InterMo representation outperforms cascade methods in generating timestamped piano sheet music from audio, even when cascades receive ground-truth MIDI.

TextTeacher: What Can Language Teach About Images?

cs.CV · 2026-05-21 · unverdicted · novelty 6.0

TextTeacher uses frozen text embeddings from captions as semantic anchors to guide vision model training, improving ImageNet accuracy by up to 2.7 p.p. and transfer performance by 1.0 p.p. on average.

Communicating Sound Through Natural Language

cs.LG · 2026-05-09 · unverdicted · novelty 6.0

Lexical acoustic coding lets LLMs transmit audio waveforms as editable natural-language sentences that another LLM can parse and reconstruct into sound.

LLM-Codec: Neural Audio Codec Meets Language Model Objectives

cs.SD · 2026-04-20 · unverdicted · novelty 6.0

LLM-Codec augments audio codec training with multi-step token prediction and contrastive semantic alignment to improve both waveform reconstruction and autoregressive predictability for speech language models.

citing papers explorer

Showing 50 of 57 citing papers.