News
우리는 다른 요원들도 배우고 있다는 사실을 고려한 알고리즘을 공개하고, 반복되는 죄수의 d에서 이기적이면서도 협력적인 전략을 발견합니다...
출처 제공 본문
We’re releasing an algorithm which accounts for the fact that other agents are learning too, and discovers self-interested yet collaborative strategies like tit-for-tat in the iterated prisoner’s dilemma. This algorithm, Learning with Opponent-Learning Awareness (LOLA), is a small step towards agents that model other minds.
댓글 0
댓글을 불러오는 중입니다.