DeepSeek V4 Pro scores 80.6% on SWE-bench, verified at $0.87 per million output tokens, versus GPT-5.6 Sol’s $30 and 91.9% Terminal-Bench score.
GPT-5.6 Sol, Claude Opus 5, and Gemini 3.1 Pro compared on benchmark scores, context window, and price per million tokens.