A 9B multimodal model learns to tailor raw video/GUI streams into schema-aligned training data, matching a proprietary annotator on downstream tasks; the abstract's capacity-scaling claims are not supported by the body.
arXiv preprint arXiv:2412.10718 (2024)
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
citation-role summary
background 1
citation-polarity summary
years
2026 2roles
background 1polarities
background 1representative citing papers
CoInteract adds a human-aware mixture-of-experts and spatially-structured co-generation to a diffusion transformer to synthesize videos with stable structures and physically plausible human-object contacts.
citing papers explorer
-
DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams
A 9B multimodal model learns to tailor raw video/GUI streams into schema-aligned training data, matching a proprietary annotator on downstream tasks; the abstract's capacity-scaling claims are not supported by the body.
-
CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation
CoInteract adds a human-aware mixture-of-experts and spatially-structured co-generation to a diffusion transformer to synthesize videos with stable structures and physically plausible human-object contacts.