Monday, August 3, 2026

AI & Models

DeepSeek launches V4 models with 1.6 trillion parameters

DeepSeek launched its V4 Flash and Pro models, which feature up to 1.6 trillion parameters and claim to almost close the performance gap with frontier AI models.

DeepSeek launches V4 models with 1.6 trillion parameters
Photo: DeepSeek

Chinese AI lab DeepSeek has launched two preview versions of its newest large language model, DeepSeek V4: DeepSeek V4 Flash and V4 Pro. The launch represents an update to the lab’s previous V3.2 model. According to DeepSeek, both are mixture-of-experts models—an AI model architecture that activates only a subset of parameters per task to lower inference costs—with context windows of 1 million tokens. DeepSeek asserts that the models have almost closed the gap with current leading models on reasoning benchmarks. However, the lab noted that the models exhibit a “developmental trajectory that trails state-of-the-art frontier models by approximately 3 to 6 months”.

The Pro model features a total of 1.6 trillion parameters, with 49 billion active parameters, making it the largest open-weight model available. This outstrips competitors such as Moonshot AI’s Kimi K 2.6, which has 1.1 trillion parameters, and MiniMax’s M1, which has 456 billion parameters. It is also more than double the parameter count of DeepSeek’s previous V3.2 model, which has 671 billion parameters. By comparison, the smaller V4 Flash model has 284 billion parameters, with 13 billion active parameters.

Both models are positioned as highly affordable alternatives to other models on the market, with pricing that undercuts major frontier models from competitors such as Google and OpenAI. According to DeepSeek’s pricing structure, the costs for the new models are as follows:

  • V4 Flash: Costs $0.14 per million input tokens and $0.28 per million output tokens.
  • V4 Pro: Costs $0.145 per million input tokens and $3.48 per million output tokens.

The launch on 2026-04-24 comes amid heightened geopolitical and industry scrutiny. The day before the launch, the U.S. accused China of stealing American AI labs’ IP on an industrial scale using thousands of proxy accounts. Furthermore, DeepSeek itself has been accused by Anthropic and OpenAI of “distilling,” essentially copying, their AI models. Distilling is defined as the process of training a smaller model using the outputs of a larger, pre-existing model.

Why it matters

DeepSeek’s new V4 models introduce a 1.6 trillion parameter open-weight model that claims to close the performance gap with frontier models while undercutting them significantly on price.