Prisoner

🧠 Asymmetric architecture. More intelligence, less cost. 🔹 552B-parameter MoE. 🔹 New Causal Encoder–Decoder architecture: just 8B active parameters for input, 16B for output. 🔹 New pre-training methods + larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro. 2/6
DeepSeek (@deepseek_ai)
🤢🤢🤢
Disgraceful benchmaxxing.
Qwen 3.8 Flash >> DeepSeek V4 Flash 0731 >> DeepSeek V4.1 Flash