ZPPO improves distillation to small vision-language models by using binary and negative candidate prompts plus a replay buffer for hard questions, outperforming standard distillation and GRPO on a 31-benchmark suite with largest gains at the 0.8B scale.
Prediction of deep ice layer thickness using adaptive recurrent graph neural networks, in: 2023 IEEE International Conference on Image Processing (ICIP), pp
4 Pith papers cite this work, alongside 26 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 4roles
baseline 1polarities
baseline 1representative citing papers
A new six-domain benchmark shows existing AI-image detectors are highly inconsistent on text-rich images and fail badly under JPEG compression, while a vision-language model is stronger but still weak on tables.
AXPO addresses the Thinking-Acting Gap in agentic RL training by targeted resampling of tool calls in all-wrong subgroups, delivering +1.8pp gains over GRPO on nine multimodal benchmarks with an 8B model beating a 32B baseline on Pass@4.
Adding MAR climate-model features plus a multi-branch GraphSAGE/temporal-convolution design with adaptive fusion reduces deep-ice-layer thickness RMSE by 21.01% over a no-knowledge multi-branch baseline on the SRED Greenland dataset.
citing papers explorer
-
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients
ZPPO improves distillation to small vision-language models by using binary and negative candidate prompts plus a replay buffer for hard questions, outperforming standard distillation and GRPO on a 31-benchmark suite with largest gains at the 0.8B scale.
-
TextRich: A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2
A new six-domain benchmark shows existing AI-image detectors are highly inconsistent on text-rich images and fail badly under JPEG compression, while a vision-language model is stronger but still weak on tables.
-
Agent Explorative Policy Optimization for Multimodal Agentic Reasoning
AXPO addresses the Thinking-Acting Gap in agentic RL training by targeted resampling of tool calls in all-wrong subgroups, delivering +1.8pp gains over GRPO on nine multimodal benchmarks with an 8B model beating a 32B baseline on Pass@4.
-
K-STEMIT: Knowledge-Informed Spatio-Temporal Efficient Multi-Branch Graph Neural Network for Subsurface Stratigraphy Thickness Estimation from Radar Data
Adding MAR climate-model features plus a multi-branch GraphSAGE/temporal-convolution design with adaptive fusion reduces deep-ice-layer thickness RMSE by 21.01% over a no-knowledge multi-branch baseline on the SRED Greenland dataset.