A recent research update on compute-efficient pretraining and scaling to trillion-parameter models reveals a significant breakthrough. The new pretraining recipe is now over 10 times more compute-efficient than leading open-weight base models. This efficiency gain allows for the matching of DeepSeek V4 Pro Base using approximately 50 times fewer FLOPs, which translates to about half of GPT3's pretraining compute or around $0.5 million on GB200.

The researchers found that by compounding algorithmic efficiency over time, they were able to achieve this substantial improvement. They continued scaling up, spending around $4 million, and meaningfully outperformed all publicly available open base models on perplexity evaluations. These results suggest that training a model of this capability would cost over $100 million under the DeepSeek V4 Pro recipe, highlighting the significance of the breakthrough.

The team believes that pretraining, agentic RL, and long-context are sufficient to build superhuman coding agents and automate AI R&D. They started by focusing on long-context and have made substantial progress in pretraining. The evaluation of the model's performance was conducted on heldout data, including code evaluations, reasoning on math problems, and text evaluations on recent research papers. The results show significant improvements across various domains.

To measure generalization, the researchers evaluated loss on heldout data and found that their model outperforms leading open models. They also tested the model's knowledge in key domains to identify gaps in the dataset. By collecting granular buckets of content, they can get more precise signals on the model's performance.

The team's goal is to build the best model for coding and autonomous AI R&D. To achieve this, they evaluate domains they deprioritize, such as facts about notable people, local news, or sports/events. The evaluation results show that their recipe is more compute-efficient than the best open models per evaluation.

The researchers found that there are no shortcuts to achieving significant improvements in pretraining efficiency. They had to build a stable foundation first, which included smooth convergence, low-precision training quality, fast and stable infrastructure, and correct hyperparameter scaling rules. They also had to hunt down bugs and evaluate each model, optimizer, or data change by training three models spanning two orders of magnitude of compute.

Looking ahead, the team plans to scale long-horizon RL, training agents to keep learning after deployment through long-context. They are also working on alignment training techniques that present robust theoretical properties. With their pretraining and long-context work now quite mature, they are poised to make further significant advancements in the field.

Este artículo fue escrito con la asistencia de IA.
News Factory APP - noticias agénticas para impulsar tu SEO y AEO.