Jeonggyu Huh, Yeoneung Kim, Seungwon Jeong · 2026-08-18
A plain-English AI summary of what this paper means for investors — generated on demand from the abstract.
We develop simulation-based policy iteration for continuous-time portfolio choice with predictable returns and convex constraints. Each outer step re-evaluates a fixed-latent OL-BPTT adjoint after deployment and solves the constrained update. Shifted-adjoint cancellation controls the adjoint--HJB Hamiltonian-gradient discrepancy by the policy-improvement residual. For CRRA portfolios, exact HJB policy iteration identifies the optimal reduced value factor, while population OL-BPTT iteration converges globally under an occupation-measure relative-error condition. A theorem-matched audit yields a maximal 95% upper endpoint of 0.074 against the required 0.75 threshold. In a three-factor, fifty-asset design, current-policy re-evaluation outperforms matched pooled refinement under both evaluation laws.
Go deeper: a full research-committee breakdown of this paper, its assumptions and failure modes, and how its method would apply to a specific ticker or your watchlist. See StockTools AI →
AI summary generated from the paper’s public abstract via arXiv; it may miss nuance — read the source before relying on it. Thank you to arXiv for its open-access interoperability; StockTools is not affiliated with arXiv, and all rights remain with the authors. Educational only, not financial advice.