News
우리는 한 요원을 훈련시켜 단일 인간 시연으로 몬테수마의 복수에서 74,500점이라는 최고 점수를 달성하도록 훈련시켰습니다. 이는 이전에 발표된 어떤 결과보다도 우수합니다. 우리의 알고리즘은 간단합니다: 에이전트 pl...
출처 제공 본문
We’ve trained an agent to achieve a high score of 74,500 on Montezuma’s Revenge from a single human demonstration, better than any previously published result. Our algorithm is simple: the agent plays a sequence of games starting from carefully chosen states from the demonstration, and learns from them by optimizing the game score using PPO, the same reinforcement learning algorithm that underpins OpenAI Five.
댓글 0
댓글을 불러오는 중입니다.