Learning-Based Min-Max Differential Dynamic Programming

Loading...
Thumbnail Image

Date

Authors

Sarbaz, Mohammad

Journal Title

Journal ISSN

Volume Title

Publisher

University of Oklahoma – Graduate College

Item Statistics

  • Total Views: 34
  • Total Downloads: 535
  • Views in the Last Month: 2

Abstract

Designing optimal control strategies for dynamic systems operating in uncertain and adversarial environments remains a fundamental challenge in modern control theory and reinforcement learning. Traditional control methods often assume precise knowledge of system dynamics, but real-world systems frequently operate with incomplete information, disturbances, and unpredictable adversarial influences. This has led to the development of robust and adaptive control strategies that can optimize performance while mitigating worst-case uncertainties. Among these, Min-Max Differential Dynamic Programming (Min-Max DDP) stands out as an effective approach, leveraging the principles of dynamic programming to iteratively refine control policies while explicitly accounting for worst-case adversarial disturbances. By framing the control problem as a minimization–maximization game, Min-Max DDP explicitly considers worst-case disturbances, providing robustness in safety-critical and adversarial settings. Despite its advantages, Min-Max DDP faces several key challenges that limit its practical applicability. However, conventional implementations rely on exact system models, limiting their applicability to systems with unknown or stochastic dynamics. Additionally, extending Min-Max DDP to large-scale, multi-agent, and output feedback systems introduces further complexity. On the other hand, ensuring safety in the framework’s performance is a crucial aspect that cannot be neglected. This research systematically addresses these challenges by developing a series of novel methodologies that enhance the adaptability, robustness, and scalability of Min-Max DDP. To overcome the limitation of requiring exact dynamics, data-driven Min-Max DDP is introduced, employing Gaussian Process Regression to approximate system dynamics and compute optimal control updates in the absence of explicit models. The framework is further extended to a stochastic game-theoretic Min-Max DDP formulation that models both drift and diffusion components. Gaussian Process Regression is used to approximate the drift, while a binning-based nonparametric method is employed to estimate the diffusion dynamics, enabling a fully data-driven framework for stochastic Min-Max control. To enhance scalability, the research develops a distributed Min-Max DDP framework enabling decentralized coordination among interconnected subsystems. This approach allows interconnected agents to collaborate effectively while maintaining robustness against uncertainty. Additionally, Min-Max DDP is applied to stochastic $H_{\infty}$ control as a robust strategy to mitigate disturbances and ensure stability. When direct state observation is unavailable, the research introduces sliding mode observer-based Min-Max DDP, integrating sliding mode observers to estimate system states and provide robust output feedback control, ensuring effective performance despite limited observability. The research further extends Min-Max DDP to nonzero-sum and multiplayer differential games, enabling optimal strategies in competitive and cooperative multi-agent settings. Additinally, the study investigates Min-Max adaptive dynamic programming for zero-sum differential games, incorporating adaptive learning mechanisms to enhance performance in adversarial settings. Finally, this study incorporates a chance-constrained framework to ensure that the trajectory optimization satisfies safety requirements. By systematically addressing these challenges, the work transforms Min-Max DDP into a versatile framework for data-driven, robust, and multi-agent control. Through the integration of machine learning, stochastic modeling, and distributed control, these contributions advance the frontiers of optimal control, making Min-Max Differential Dynamic Programming more applicable to real-world problems characterized by uncertainty, partial observability, and adversarial interactions. The proposed methods are validated on diverse dynamic systems, including the inverted pendulum, double inverted pendulum, mass–spring–damper, power systems, two-degree-of-freedom helicopter, and nonlinear Quadrotor. Additional tests on vertical flight regulation, cart-pole, unicycle robot, and load-frequency control further demonstrate their effectiveness across varied control scenarios.

Description

Citation

Related file

Notes

Endorsement

Review

Supplemented By

Referenced By

DOI

Collection Detail

# of Isolates from RBM

# of Isolates from TV8