Qwen3.7-Max: Alibaba Ships a Model That Runs for 35 Hours Straight
Alibaba's new flagship tops the Chinese leaderboard, plugs into Claude Code, and claims day-long autonomous runs — a clear shot at the agentic frontier.
- 01Qwen3.7-Max launched 19–20 May 2026 at Alibaba's Cloud Summit, entering the Artificial Analysis Intelligence Index at 56.6 — the highest-ranked Chinese model to date.
- 02Alibaba claims it can run autonomously for up to 35 hours and supports external harnesses like Anthropic's Claude Code, underscoring an agentic, long-horizon focus.
- 03Qwen3.7-Max is proprietary and API-only via Alibaba Cloud Model Studio, with a 1M-token context — a contrast to Qwen's open-weight heritage.

Alibaba's Qwen team built its reputation on open weights — permissively licensed models that anyone could download. Its newest flagship, Qwen3.7-Max, launched around 19–20 May 2026 at the Alibaba Cloud Summit, breaks from that mould in one important way, and pushes hard on the frontier in another.
Top of the Chinese leaderboard
Qwen3.7-Max entered the Artificial Analysis Intelligence Index at 56.6 — the highest-ranked Chinese AI model to date on that leaderboard. The supporting benchmarks are strong: 92.4 on GPQA Diamond, 69.7 on Terminal-Bench 2.0-Terminus, and competitive scores on Humanity's Last Exam and the MCP-Atlas coding-agent benchmark. On Alibaba's own evaluations, it wins or ties against comparable frontier models on a range of reasoning, long-context and multilingual tasks. (As always, treat lab-reported wins against named competitors as directional rather than gospel — independent evals are the better guide.)
The 35-hour claim
The most eye-catching figure isn't a benchmark. Alibaba says Qwen3.7-Max can run autonomously for up to 35 hours on a task, and that it supports external harnesses like Anthropic's Claude Code. Whether or not 35 hours is achievable on real-world work, the framing is the point: Alibaba is explicitly competing on long-horizon autonomy — how long a model can grind on a complex job without losing the thread — which is exactly the axis Anthropic and Google are racing on. The Claude Code interoperability is a notably pragmatic touch: rather than forcing its own tooling, Alibaba is meeting developers in the agent harnesses they already use.
The proprietary turn
Here's the break from Qwen's heritage: Qwen3.7-Max is proprietary and API-only, served via Alibaba Cloud Model Studio, with a 1-million-token context window. There's no open download of the Max tier. That's a meaningful strategic shift for a team whose brand is open weights — and it mirrors a broader pattern where labs keep their absolute-frontier model closed (and monetised via API) while open-sourcing smaller tiers. Qwen's open lineage continues lower down the stack; the very top is now behind an API.
Why it matters
The old story about Chinese labs was 'fast follower, cheaper, a bit behind.' Qwen3.7-Max complicates that. A model topping the Chinese leaderboard, claiming day-long autonomous runs, and interoperating with Western agent tooling is competing on the frontier's own terms — autonomy and agentic reliability — not just on price. For anyone building agentic systems, it's a credible non-Western option, with the obvious caveat that it's API-only and runs on Alibaba's cloud.
Ask Relay — he reads every question himself and replies personally by email.
