MemLearner introduces a learning-based adaptive context query method using query tokens in video world models to improve long-term scene consistency over rule-based retrieval.
hub
Advances in neural information processing systems33, 1877–1901 (2020)
12 Pith papers cite this work. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
years
2026 12roles
background 1polarities
background 1representative citing papers
PromptDx adds a differentiable adapter to align multimodal data with a pre-trained TabPFN-style ICL engine, achieving strong Alzheimer's diagnosis performance with only 1% context samples.
SceneOrchestra trains an orchestrator to generate full tool-call trajectories for 3D scene synthesis and uses a discriminator during training to select high-quality plans, yielding state-of-the-art results with lower runtime.
CALIBER conditions the variational posterior of low-rank adapters on token-level cross-attention between text and audio to produce uncertainty-aware multimodal parameter-efficient fine-tuning.
Phi-Nav generates path-level hindsight instructions from on-policy exploration trajectories to supply additional semantic supervision for vision-language navigation agents.
Minos uses a two-tiered multi-agent architecture with retrieval-augmented reasoning and FSM-coordinated agents to reconstruct attack scenarios from provenance data, reporting 0.92 recall and 0.64 precision on 14 scenarios.
RosettaSim adapts frozen LLMs via structured autoregressive modeling of scene topology and agent states to reach SOTA short- and long-term traffic simulation on WOSAC, paired with RTE evaluation that correlates better with human-like fidelity.
Introduces CORTEX benchmark supplying 76,177 validated four-stage diagnostic reasoning traces for open/closed VQA and report generation on chest CT to enable traceable MLLM supervision and evaluation.
A sequential-to-global SSL method based on DINO pretrains iterative foveal-inspired vision transformers to achieve competitive ImageNet-1K performance with constant compute regardless of input resolution.
LiteSemRAG delivers leading MRR@10 on three benchmarks using only lightweight semantic graph methods and zero LLM tokens.
An asynchronous fast-slow dual-system with DiT action modeling and time-weighted loss doubles unseen aerial VLN success rates and halves decision latency in simulation.
citing papers explorer
-
MemLearner: Learning to Query Context memory for Video World Models
MemLearner introduces a learning-based adaptive context query method using query tokens in video world models to improve long-term scene consistency over rule-based retrieval.
-
PromptDx: Differentiable Prompt Tuning for Multimodal In-Context Alzheimer's Diagnosis
PromptDx adds a differentiable adapter to align multimodal data with a pre-trained TabPFN-style ICL engine, achieving strong Alzheimer's diagnosis performance with only 1% context samples.
-
SceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation
SceneOrchestra trains an orchestrator to generate full tool-call trajectories for 3D scene synthesis and uses a discriminator during training to select high-quality plans, yielding state-of-the-art results with lower runtime.
-
Cross-Modal Bayesian Low-Rank Adaptation for Uncertainty-Aware Multimodal Learning
CALIBER conditions the variational posterior of low-rank adapters on token-level cross-attention between text and audio to produce uncertainty-aware multimodal parameter-efficient fine-tuning.
-
Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation
Phi-Nav generates path-level hindsight instructions from on-policy exploration trajectories to supply additional semantic supervision for vision-language navigation agents.
-
Minos: A Multi-Agent Collaborative Framework for Provenance-Based Backward Tracking
Minos uses a two-tiered multi-agent architecture with retrieval-augmented reasoning and FSM-coordinated agents to reconstruct attack scenarios from provenance data, reporting 0.92 recall and 0.64 precision on 14 scenarios.
-
Long-term Traffic Simulation via Structured Autoregressive Modeling
RosettaSim adapts frozen LLMs via structured autoregressive modeling of scene topology and agent states to reach SOTA short- and long-term traffic simulation on WOSAC, paired with RTE evaluation that correlates better with human-like fidelity.
-
CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs
Introduces CORTEX benchmark supplying 76,177 validated four-stage diagnostic reasoning traces for open/closed VQA and report generation on chest CT to enable traceable MLLM supervision and evaluation.
-
Self-supervised pretraining for an iterative image size agnostic vision transformer
A sequential-to-global SSL method based on DINO pretrains iterative foveal-inspired vision transformers to achieve competitive ImageNet-1K performance with constant compute regardless of input resolution.
-
LiteSemRAG: Lightweight LLM-Free Semantic-Aware Graph Retrieval for Robust RAG
LiteSemRAG delivers leading MRR@10 on three benchmarks using only lightweight semantic graph methods and zero LLM tokens.
-
FSD-VLN: Fast-Slow Dual-System Modeling for Aerial Long-Horizon Vision-Language Navigation
An asynchronous fast-slow dual-system with DiT action modeling and time-weighted loss doubles unseen aerial VLN success rates and halves decision latency in simulation.
- NaLA: A 3D Native LLM Layout Agent for High-quality 3D Scene Generation