Submitted by JGC 37 PVPO: Pre-Estimated Value-Based Policy Optimization for Agentic Reasoning Apsara Stack MaaS 2