Loki replaces RGB conditioning stacks with identity-orthogonal parametric face encodings rasterized for diffusion, achieving efficient cross-ID portrait animation without cross-ID training data.
hub
Film: Visual reasoning with a general conditioning layer
14 Pith papers cite this work. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
roles
method 1polarities
use method 1representative citing papers
AffectCodec applies block-diagonal projections in residual FSQ to explicitly allocate bits to emotion and acoustic subspaces, combined with emotion conditioning, yielding better emotion preservation at low bitrates with competitive acoustic quality.
A hypernetwork generates complete task-specific visuomotor policy parameters from instructions alone to structurally eliminate observation leakage in language-conditioned robotic control.
DualLGD reformulates molecular graph denoising as alternating atom and bond subproblems in separate streams, achieving 34.37% and 23.89% top-1 accuracy on NPLIB1 and MassSpecGym benchmarks, roughly 3x prior state of the art.
CoExVQA uses a chain-of-explanation to ground DocVQA answers in localized document regions, achieving state-of-the-art explainable performance with a 12% ANLS gain on PFL-DocVQA over prior baselines.
SpecSem-Net integrates Fourier-based spectral filtering with semantic-guided gated merging to detect AI-generated videos, reporting 87.25% accuracy on a new benchmark of five commercial generators and 95.59% on public datasets.
XDecomposer uses set prediction and phase-query decomposition to jointly identify phases and reconstruct multiphase PXRD patterns without priors.
ViewSAM achieves state-of-the-art weakly supervised performance on cross-view referring multi-object tracking by refining SAM tracklets via affinity-guided re-prompting and modeling view-induced variations as learnable conditions on SAM2.
A fine-tuned video diffusion model becomes a fast, differentiable CFD surrogate for urban wind, enabling gradient-based building-layout optimization confirmed by ground-truth simulations.
Reinforcement learning with a tunable control parameter and clinical reward enables precision-recall controllable radiology report generation that outperforms prior methods on MIMIC-CXR.
A continuous data assimilation framework enables a-priori training of solver-conditioned neural turbulence closures that remain stable at deployment and track discretization errors.
One physics-informed network trained on the heat equation, not labeled data, predicts laser-scan temperature fields for unseen alloys including copper with ~1% relative error.
AVA-VLA reformulates VLA learning as a POMDP using recurrent states and active visual attention to achieve state-of-the-art results on LIBERO, CALVIN, and real dual-arm tasks.
GR-3 is a VLA model that generalizes to novel objects, environments, and abstract instructions, outperforms the π0 baseline, and integrates with the new ByteMini bi-manual mobile robot.
citing papers explorer
-
Loki: Representation over Architecture for Diffusion-Based Portrait Animation
Loki replaces RGB conditioning stacks with identity-orthogonal parametric face encodings rasterized for diffusion, achieving efficient cross-ID portrait animation without cross-ID training data.
-
AffectCodec: Emotion-Preserving Neural Speech Codec with Block-Diagonal Residual FSQ
AffectCodec applies block-diagonal projections in residual FSQ to explicitly allocate bits to emotion and acoustic subspaces, combined with emotion conditioning, yielding better emotion preservation at low bitrates with competitive acoustic quality.
-
DISC: Decoupling Instruction from State-Conditioned Control via Policy Generation
A hypernetwork generates complete task-specific visuomotor policy parameters from instructions alone to structurally eliminate observation leakage in language-conditioned robotic control.
-
Unlocking High-Fidelity Molecular Generation from Mass Spectra via Dual-Stream Line Graph Diffusion
DualLGD reformulates molecular graph denoising as alternating atom and bond subproblems in separate streams, achieving 34.37% and 23.89% top-1 accuracy on NPLIB1 and MassSpecGym benchmarks, roughly 3x prior state of the art.
-
Towards Self-Explainable Document Visual Question Answering with Chain-of-Explanation Predictions
CoExVQA uses a chain-of-explanation to ground DocVQA answers in localized document regions, achieving state-of-the-art explainable performance with a 12% ANLS gain on PFL-DocVQA over prior baselines.
-
SpecSem-Net: Integrating Spectral and Semantic Features for Robust AI-generated Video Detection
SpecSem-Net integrates Fourier-based spectral filtering with semantic-guided gated merging to detect AI-generated videos, reporting 87.25% accuracy on a new benchmark of five commercial generators and 95.59% on public datasets.
-
XDecomposer: Learning Prior-Free Set Decomposition for Multiphase X-ray Diffraction
XDecomposer uses set prediction and phase-query decomposition to jointly identify phases and reconstruct multiphase PXRD patterns without priors.
-
ViewSAM: Learning View-aware Cross-modal Semantics for Weakly Supervised Cross-view Referring Multi-Object Tracking
ViewSAM achieves state-of-the-art weakly supervised performance on cross-view referring multi-object tracking by refining SAM tracklets via affinity-guided re-prompting and modeling view-induced variations as learnable conditions on SAM2.
-
Pretrained Video Models as Differentiable Physics Simulators for Urban Wind Flows
A fine-tuned video diffusion model becomes a fast, differentiable CFD surrogate for urban wind, enabling gradient-based building-layout optimization confirmed by ground-truth simulations.
-
Precision Recall Controllable Radiology Report Generation via Hybrid Natural Language and Clinical Reward Learning
Reinforcement learning with a tunable control parameter and clinical reward enables precision-recall controllable radiology report generation that outperforms prior methods on MIMIC-CXR.
-
Deep Learning of Solver-Aware Turbulence Closures from Nudged LES Dynamics
A continuous data assimilation framework enables a-priori training of solver-conditioned neural turbulence closures that remain stable at deployment and track discretization errors.
-
Material-agnostic temperature field prediction for metal additive manufacturing via a parametric PINN framework
One physics-informed network trained on the heat equation, not labeled data, predicts laser-scan temperature fields for unseen alloys including copper with ~1% relative error.
-
AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
AVA-VLA reformulates VLA learning as a POMDP using recurrent states and active visual attention to achieve state-of-the-art results on LIBERO, CALVIN, and real dual-arm tasks.
-
GR-3 Technical Report
GR-3 is a VLA model that generalizes to novel objects, environments, and abstract instructions, outperforms the π0 baseline, and integrates with the new ByteMini bi-manual mobile robot.