Reproducing DeepSeek R1: Can You Run RL on 15GB VRAM?

CategoryNews Briefs

Recent attempts to reproduce DeepSeek R1’s open-source work have been everywhere. I went through the paper and community feedback, and a few counterintuitive findings are worth discussing.

1. 15 GB of VRAM = the new entry-level bar?

The Unsloth team pushed optimization to the limit, training a 15-billion-parameter model on just 15 GB of VRAM. What does that mean? You can run GRPO (Group Relative Policy Optimization) reinforcement learning fine-tuning on Google Colab’s free tier. What used to feel like a playground reserved for big tech companies is now accessible to individual developers.

2. The “aha moment” may be a myth

Research from Sea AI Lab drops some cold water on this: R1-style “self-reflection” and “insight” aren’t actually produced by RL. They’re a form of shallow self-reflection baked in during pretraining. RL merely optimizes the reward function, coaxing longer answers out of the model rather than genuine reasoning. Don’t fall for the marketing—small models don’t spontaneously generate big intelligence.

3. Data quality beats quantity

The LIMO dataset uses only 817 hand-picked samples yet outperforms far larger models on AIME and MATH benchmarks. Meanwhile, Qwen2.5-32B fine-tuned on the s1K dataset (1,000 math problems) scores competitively with OpenAI o1-preview. What’s the takeaway? Carefully curated “diamonds” in your data matter far more than a mountain of “gravel.”

4. New patterns for industrial adoption

Beyond hyperparameter tuning, Microsoft’s PIKE-RAG framework is worth a closer look. It builds multi-layer heterogeneous graphs to extract private-domain knowledge, specifically targeting the gap where LLMs stumble over industry jargon. Paired with DeepRAG’s “on-demand retrieval” mechanism—which cuts irrelevant retrievals by 47% and lifts accuracy by 22%—this finally gives us a replicable path for building local knowledge bases.

The bottom line: The future belongs to teams that master fine-grained operations around “data curation + low-cost fine-tuning,” not those that simply stack compute. If you’ve got a good GPU, you really can get started.

Source · Moresso Notes: Read original →

Get the Creator Daily by email
Hand-picked opportunities, tools & insights for indie makers — free.
中文读者?订阅中文频道 →
iMessage 邮件 Contact us
中文