GEOZ

Entity

DeepSeek-R1-Zero

A variant of DeepSeek-R1 trained directly with reinforcement learning on the base model, skipping traditional supervised fine-tuning (SFT).

术语属性

  • Training MethodReinforcement Learning without SFT
  • Performance71.0% pass@1 on AIME 2024
  • Voting Performance86.7% with majority voting

相关文章