MAVIC corrects Bellman backups at instruction boundaries by adjusting the incoming objective and restoring continuation value, enabling consistent estimation under stochastic instruction switching in cooperative MARL.
Title resolution pending
2 Pith papers cite this work, alongside 10 external citations. Polarity classification is still indexing.
2
Pith papers citing it
10
external citations · external index
citation-role summary
background 1
citation-polarity summary
fields
cs.AI 2years
2026 2verdicts
UNVERDICTED 2roles
background 1polarities
background 1representative citing papers
A 4B-parameter local LLM trained with tool-augmented process-rewarded learning generates STL formulas from natural language at state-of-the-art accuracy on a new bilingual benchmark.
citing papers explorer
-
Robust Instruction Compliance in Cooperative Multi-Agent Reinforcement Learning
MAVIC corrects Bellman backups at instruction boundaries by adjusting the incoming objective and restoring continuation value, enabling consistent estimation under stochastic instruction switching in cooperative MARL.
-
ReasonSTL: Bridging Natural Language and Signal Temporal Logic via Tool-Augmented Process-Rewarded Learning
A 4B-parameter local LLM trained with tool-augmented process-rewarded learning generates STL formulas from natural language at state-of-the-art accuracy on a new bilingual benchmark.