ClawArena introduces a benchmark with hidden ground truth, noisy multi-channel traces, and 14-category questions to evaluate multi-source conflict reasoning, dynamic belief revision, and implicit personalization in AI agents.
Title resolution pending
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3verdicts
UNVERDICTED 3representative citing papers
OmniManim improves render quality in educational animation code generation by using a Vision Agent with coarse-to-fine bounding-box denoising and interpolation-aware optimization on new datasets.
CAGE uses LLM-generated code for label-correct diagrams followed by ControlNet-conditioned diffusion refinement to produce both accurate and visually engaging educational graphics, backed by the new EduDiagram-2K dataset.
citing papers explorer
-
ClawArena: Benchmarking AI Agents in Evolving Information Environments
ClawArena introduces a benchmark with hidden ground truth, noisy multi-channel traces, and 14-category questions to evaluate multi-source conflict reasoning, dynamic belief revision, and implicit personalization in AI agents.
-
See Before You Code: Learning Visual Priors for Spatially Aware Educational Animation Generation
OmniManim improves render quality in educational animation code generation by using a Vision Agent with coarse-to-fine bounding-box denoising and interpolation-aware optimization on new datasets.
-
CAGE: Bridging the Accuracy-Aesthetics Gap in Educational Diagrams via Code-Anchored Generative Enhancement
CAGE uses LLM-generated code for label-correct diagrams followed by ControlNet-conditioned diffusion refinement to produce both accurate and visually engaging educational graphics, backed by the new EduDiagram-2K dataset.