Qwen 3.7 Max Benchmarks: Performance, Model Comparisons and What the Results Really Show
Qwen 3.7 Max posts some of Qwen’s strongest benchmark results yet, scoring 80.4 on SWE-bench Verified, 60.6 on SWE-Pro, 92.4 on GPQA Diamond and 91.6 on LiveCodeBench. It leads or closely challenges models including Opus 4.6 Max, DeepSeek V4 Pro Max, GLM-5.1 Thinking and Qwen 3.6 Plus across several coding, reasoning and agent evaluations. The … Read more