HPML projects multi-agent update fields onto the closest metric-gradient potential flow via Hodge decomposition, yielding Lyapunov potentials and equilibrium-gap bounds.
Lisfc-search: Lifelong search for network sfc optimization under non-stationary drifts.arXiv preprint arXiv:2602.14360, 2026a
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3representative citing papers
VPSD-RL discovers exact and approximate value-preserving Lie-group operators in continuous RL to stabilize learning via transition augmentation and consistency regularization.
A 13M-parameter intent-conditioned reward model trained on 398K multi-OS GUI steps scores candidate actions and lifts Agent S3 OSWorld success by 6.9 points without extra LLM calls.
citing papers explorer
-
Metric-Gradient Projection for Stable Multi-Agent Policy Learning
HPML projects multi-agent update fields onto the closest metric-gradient potential flow via Hodge decomposition, yielding Lyapunov potentials and equilibrium-gap bounds.
-
Operator-Guided Invariance Learning for Continuous Reinforcement Learning
VPSD-RL discovers exact and approximate value-preserving Lie-group operators in continuous RL to stabilize learning via transition augmentation and consistency regularization.
-
IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents
A 13M-parameter intent-conditioned reward model trained on 398K multi-OS GUI steps scores candidate actions and lifts Agent S3 OSWorld success by 6.9 points without extra LLM calls.