The Open-Weight Scoreboard, Mid-2026: Cohere, Mistral, DeepSeek and Qwen
Permissive licences, two-GPU footprints, trillion-parameter MoEs — the downloadable tier has turned 'open' from a compromise into a deployment strategy.
- 01Open weights had a strong run into mid-2026: Cohere's Apache 2.0 Command A+ (May), DeepSeek's V4 Preview (April), and Mistral's Apache 2.0 Large 3 (Dec 2025) anchor the tier.
- 02The competition among open models is now on deployability — licence terms, GPU footprint, context length, languages — not just raw capability.
- 03The Max-tier paradox: even open-heritage labs (Qwen) now keep their very top model closed and API-only, while open-sourcing everything below it.

Two years ago, choosing an open-weight model meant accepting a visible quality drop in exchange for control or cost. By mid-2026 that trade has largely inverted for everyday production work — and the releases of the last few months show why. Here's the open-weight scoreboard.
Cohere Command A+ — the enterprise-deployability play
Released 20 May 2026, Command A+ is a 218B sparse MoE (25B active) and Cohere's first fully Apache 2.0 model. Its standout traits aren't benchmark records — they're deployment ones: near-lossless quantization lets it run on two H100s or a single Blackwell B200, with a 128K context, 48 languages, and native citation grounding. It's engineered for sovereign and on-prem enterprise use, folding four prior Cohere models into one.
DeepSeek V4 Preview — scale at aggressive cost
DeepSeek's V4 Preview (24 April 2026) brings raw scale: V4-Pro at 1.6T total / 49B active and V4-Flash at 284B / 13B active, both with a 1M-token context. DeepSeek's role is to keep the price floor low — every capable model it ships compresses what closed labs can charge. (Note: the much-rumoured R2 still has no official release — don't believe the 'R2 is out' posts until DeepSeek lists it.)
Mistral Large 3 — the European Apache flagship
Mistral Large 3 (released December 2025, still the 2026 flagship) is a 675B-total / 41B-active MoE under Apache 2.0 — one of the largest open-weight models from a major lab — backed by the smaller Ministral 3 dense family (14B/8B/3B, also Apache 2.0). Mistral has rounded out 2026 with the Voxtral TTS audio model (March) and the Forge enterprise training platform, keeping a credible European open-weight option on the board.
Qwen — open heritage, closed peak
Alibaba's Qwen line has long been an open-weight staple (Qwen3, Apache 2.0). But its newest flagship, Qwen3.7-Max (May), is proprietary and API-only. That's the open-weight paradox of mid-2026 in one model: labs open-source the tiers that build community and adoption, then keep the absolute frontier closed and monetised. OpenAI did the same in reverse — shipping open gpt-oss reasoning models (120b/20b) for self-hosting while its frontier stays closed.
What the scoreboard tells you
Three patterns stand out:
- The decision moved up the stack. Among open models, you're no longer choosing on raw capability — you're choosing on licence (Apache 2.0 vs conditional), GPU footprint, context length, and language coverage.
- MoE won. Almost every serious open release is a Mixture-of-Experts — huge total capacity, modest active cost. It's how the open tier delivers frontier-adjacent quality at runnable cost.
- Open and closed coexist within each lab. The cleanest dividing line isn't 'open labs vs closed labs' anymore — it's 'open tiers vs the closed peak', often inside the same company.
For a team deciding what to run today, the honest summary is: a strong open MoE handles the bulk of real production traffic at a fraction of frontier API cost, and the remaining gap is concentrated in the hardest long-horizon agentic work. Pick your model on where your data has to live and what you can afford to run — not on a leaderboard.
Ask Relay — he reads every question himself and replies personally by email.
