MiniMax M3: Sparse Attention, 1M Context, and the Agent Model Nobody Should Evaluate Lazily

MiniMax M3 feature image showing sparse attention and 1M context evaluation

MiniMax M3 arrives with the kind of launch copy that makes engineers both curious and allergic. Frontier coding. One million tokens. Native multimodality. Open weights coming soon. Cheaper than the usual suspects. Somewhere, a product manager is already updating a roadmap slide with fireworks. The interesting part is not the fireworks. It’s the engineering bet … Read more

Qwen 3.7 Max Benchmarks: Performance, Model Comparisons and What the Results Really Show

Qwen 3.7 Max review feature image showing long-running agent engineering workflow

Qwen 3.7 Max posts some of Qwen’s strongest benchmark results yet, scoring 80.4 on SWE-bench Verified, 60.6 on SWE-Pro, 92.4 on GPQA Diamond and 91.6 on LiveCodeBench. It leads or closely challenges models including Opus 4.6 Max, DeepSeek V4 Pro Max, GLM-5.1 Thinking and Qwen 3.6 Plus across several coding, reasoning and agent evaluations. The … Read more