mimile
Back to feed

Alibaba Qwen 3.8 Flash-Next Undercuts Rivals at $0.16 per Million Tokens

AI digest

This digest was compiled by AI from multiple sources — links to the originals are below.

Alibaba Qwen 3.8 Flash-Next Undercuts Rivals at $0.16 per Million Tokens

Alibaba launched Qwen 3.8 Flash-Next, a 125B-parameter multimodal model priced at $0.16 per million input tokens and $0.47 per million output tokens. The open-weight model targets enterprise workloads with agentic capabilities and lower inference costs. Analysts caution that price alone should not drive model selection, citing data residency and security considerations.

Key Facts

  • Qwen 3.8 Flash-Next is a 125B-parameter multimodal model with 6B active parameters per token using a mixture-of-experts architecture.
  • Alibaba priced the model at $0.16 per million input tokens and $0.47 per million output tokens.
  • The model is open weight, unlike frontier models from OpenAI and Anthropic.
  • Gartner analyst Arun Chandrasekaran said enterprises should consider data residency and security implications before adopting the model.

Model Specifications

Qwen 3.8 Flash-Next is a 125B-parameter multimodal model with an active 6B parameters per token and a mixture-of-experts architecture. In comparison, Qwen 3.8 Max has 2.4 trillion parameters and Qwen 3.8-27B has 27B parameters. Rival Chinese AI vendor Moonshot's Kimi K3 MoE model has 2.8 trillion parameters. Alibaba highlighted that compared to Qwen 3.7-Plus, the Flash-Next version has lower training and inference costs.

Enterprise Positioning

The model excels at computer use and can interact with complex APIs, calculators and custom external databases using visual and text prompts, according to the vendor. Gartner analyst Arun Chandrasekaran said Alibaba can keep pricing competitive because of the model's inference efficiency, with only a few parameters active at a time. Chandrasekaran noted the model is optimized for agentic workloads such as tool calling and coding applications. He added that enterprises should examine whether to self-host the open-weight model rather than consume it via API.

Security Considerations

Chandrasekaran said enterprises should consider the model's data residency, given Chinese AI vendors' links to the Chinese government, and the associated security implications. He stated that enterprises want a cost-efficient model but not at the price of security, data residency and continuous innovation. Low price alone should not be a parameter, Chandrasekaran concluded.

1 source

Alibaba Qwen 3.8 Flash-Next Undercuts Rivals at $0.16 per Million Tokens