News
우리는 7억 7,400만 개의 GPT-2 언어 모델을 인간의 피드백을 활용해 다양한 작업을 미세 조정했으며, 외부 인간 라벨러의 선호도를 성공적으로 일치시켰습니다. 다만 그 선호도들은 n...
출처 제공 본문
We’ve fine-tuned the 774M parameter GPT-2 language model using human feedback for various tasks, successfully matching the preferences of the external human labelers, though those preferences did not always match our own. Specifically, for summarization tasks the labelers preferred sentences copied wholesale from the input (we’d only asked them to ensure accuracy), so our models learned to copy. Summarization required 60k human labels; simpler tasks which continue text in various styles required only 5k. Our motivation is to move safety techniques closer to the general task of “machines talking to humans,” which we believe is key to extracting information about human values.
댓글 0
댓글을 불러오는 중입니다.