Skip to content

Qwen 3.8 Max + Qwen 3.7 Flash — Frontier Depth Or Multimodal Speed

August 4, 2026

Two new Qwen models are available through Alibaba Model Studio's EU endpoint. Pick Qwen 3.8 Max for demanding professional and agentic work, or Qwen 3.7 Flash for fast, low-cost multimodal tasks with the same long context window.

What you can do

  • Choose Qwen 3.8 Max for frontier work — Alibaba's 2.4-trillion-parameter MoE flagship is built for long-horizon coding, professional analysis, tool use, and multi-step execution.
  • Choose Qwen 3.7 Flash for speed and value — the newer Flash model targets fast multimodal understanding, agent execution, OCR, and live coding workflows.
  • Work across text, images, and video — both models accept all three input types and return text.
  • Keep large jobs together — both models support a 1M-token context window, up to 991K input tokens, and up to 128K output tokens.
  • Use deep thinking when the task needs it — both models support reasoning; Qwen 3.8 Max can use up to 256K chain-of-thought tokens.
  • Run tools reliably — both models support function calling and structured output.
  • Use Qwen 3.8 Max in Swarm — it is available for parallel multi-agent plans; Qwen 3.7 Flash remains a single-model option.
  • Keep data on the EU route — Alfrada OS calls Alibaba Model Studio directly through its Germany (Frankfurt) endpoint instead of routing these models through OpenRouter.

Where this shows up

  • You want a model to inspect screenshots or video alongside a long technical brief and then execute a multi-step plan. Pick Qwen 3.8 Max.
  • You need quick OCR, visual analysis, coding assistance, or high-volume tool calls without giving up a 1M-token context window. Pick Qwen 3.7 Flash.
  • You are assembling a Swarm for a difficult engineering or research job and want Alibaba's strongest Qwen model among the workers. Add Qwen 3.8 Max.

Try it

  • "Review these screenshots and the attached specification, identify every inconsistency, then produce a prioritized implementation plan."
  • Pick Qwen 3.7 Flash in the picker, then: "Inspect this video and extract the decisions, owners, deadlines, and unresolved questions."
  • "Build a Swarm with Qwen 3.8 Max to audit this codebase, challenge the proposed architecture, and verify the final recommendation."

Heads up

  • Both models are Auto-eligible — Qwen 3.7 Flash joins the budget Auto pool; Qwen 3.8 Max joins the SOTA Auto pool and remains selectable for Swarm.
  • Qwen 3.6 Flash and Qwen 3.7 Max are retired from the picker. Pinned sessions roll forward to Qwen 3.7 Flash and Qwen 3.8 Max respectively.
  • Token rates differ — Qwen 3.7 Flash uses Alfrada's Light 2× token band; Qwen 3.8 Max uses the Premium 6× band.
  • Text output only — images and video are accepted as input, but both models return text rather than generated media.

Built for Alfrada OS.