Entity
DeepSeek-R1-Zero
A variant of DeepSeek-R1 trained directly with reinforcement learning on the base model, skipping traditional supervised fine-tuning (SFT).
术语属性
- Training Method:Reinforcement Learning without SFT
- Performance:71.0% pass@1 on AIME 2024
- Voting Performance:86.7% with majority voting