News
우리는 비디오 데이터를 기반으로 한 대규모 생성 모델 훈련을 탐구합니다. 구체적으로, 우리는 다양한 길이, 해상도, 그리고 다양한 해상도의 동영상과 이미지에 대해 텍스트 조건부 확산 모델을 공동으로 훈련시킵니다...
출처 제공 본문
We explore large-scale training of generative models on video data. Specifically, we train text-conditional diffusion models jointly on videos and images of variable durations, resolutions and aspect ratios. We leverage a transformer architecture that operates on spacetime patches of video and image latent codes. Our largest model, Sora, is capable of generating a minute of high fidelity video. Our results suggest that scaling video generation models is a promising path towards building general purpose simulators of the physical world.
댓글 0
댓글을 불러오는 중입니다.