News
우리는 단순히 올바른 f... 질문을 보상하는 대신 각 올바른 추론 단계("프로세스 감독")를 보상함으로써 수학적 문제 해결의 새로운 최첨단 모델을 훈련시켰습니다.
출처 제공 본문
We’ve trained a model to achieve a new state-of-the-art in mathematical problem solving by rewarding each correct step of reasoning (“process supervision”) instead of simply rewarding the correct final answer (“outcome supervision”). In addition to boosting performance relative to outcome supervision, process supervision also has an important alignment benefit: it directly trains the model to produce a chain-of-thought that is endorsed by humans.
댓글 0
댓글을 불러오는 중입니다.