DeepSeek is testing a sparse attention technique aimed at cutting the processing costs of large AI language models. By limiting the number of word‑to‑word comparisons, the approach seeks to mitigate the quadratic scaling problem inherent in traditional transformer architectures. The effort could make long‑form interactions more affordable while maintaining the model’s ability to understand context.
Lire la suite