News
우리는 최신 접근법과 동등하거나 더 나은 성능을 보여주면서도, 훨씬 더 간단하게 구현할 수 있는 새로운 강화 학습 알고리즘 클래스인 근접 정책 최적화(PPO)를 출시합니다.
출처 제공 본문
We’re releasing a new class of reinforcement learning algorithms, Proximal Policy Optimization (PPO), which perform comparably or better than state-of-the-art approaches while being much simpler to implement and tune. PPO has become the default reinforcement learning algorithm at OpenAI because of its ease of use and good performance.
댓글 0
댓글을 불러오는 중입니다.