Qwen-Audio-3.1 released with TTS-Next, ASR-Next, and price cuts up to 95%
Alibaba Apsara 2026: New Chip, RSI Plans, and Wave of Model Releases
Alibaba is releasing hardware, frontier model plans, and multiple consumer-facing AI products in a compressed window, spanning chips, mobile agents, audio, image generation, and real-time translation. The V900 chip and RSI plans position Alibaba as a vertically integrated competitor in AI infrastructure at a time when Nvidia's market presence in China is constrained.
The full picture
At Apsara Conference 2026, Alibaba announced a cluster of hardware and software developments. CEO Eddie Wu stated that machine-generated thinking is currently less than 3% of total human cognitive volume and set a goal of 1,000x human cognitive capacity. The Qwen team is pursuing Recursive Self-Improvement and planning a 5-10 trillion parameter model targeting complex, longer-horizon tasks on the path to ASI. On hardware, T-Head introduced the Zhenwu V900 chip, claiming roughly 3x the performance of its predecessor the M890, 216GB of on-package memory, 1.2 TB/s chip-to-chip interconnect, native FP8/FP4 support, and support for clusters of up to 500,000 cards for training and inference. Mass production and commercial release are targeted for Q1 2027; a successor chip, the J900, is on the roadmap for Q3 2028. All V900 performance figures are vendor-reported and have not been independently benchmarked. The predecessor M890, released May 2026, has shipped over 560,000 units to more than 400 external customers across 20 industries. Alibaba Cloud targets exceeding 20GW of global data-center capacity by 2032.
On the product side, the Qwen team launched Qwen Intelligence, a mobile AI platform with three agents. The Mobile Planner Agent is claimed to rank first on MobilePA-Bench, MobilePA-Bench Business, and Memory benchmarks. The Mobile-Use Agent scores 82.1 on MobileWorld, 92.2 on MobileWorld-Real, 97.2 on AndroidDaily, and a 90% end-to-end success rate. The Mobile Creative Agent generates images in 3 seconds, described as about 2x faster than leading competitors. Qwen is open-sourcing a benchmark suite covering planning, cross-app execution, real-device performance, and safety. Separately, reporting from late 2025 noted Alibaba deploying over 100 engineers to embed Qwen as a task-executing agent in Taobao and other retail apps, with a longer-term plan to monetize AI assistance in mobile shopping.
Qwen-Audio-3.1 upgrades existing ASR, TTS, and Realtime models and adds TTS-Next and ASR-Next. TTS-Next uses a unified language model plus diffusion framework to generate voice, sound effects, and background audio in a single pass. ASR-Next supports multi-speaker recognition with speaker labels, timestamps, emotion detection, and audio reasoning. Realtime mode supports simultaneous speaking and listening with interruption support and empathetic responses based on detected mood. Price cuts accompany the release: TTS by roughly 70%, Realtime by roughly 85%, and ASR by up to 95%.
Qwen-Image-2.1, a 7B open-weight model handling both image generation and editing from a single checkpoint, was released with a live demo on Hugging Face Spaces. The model supports up to 10 reference images, natively generates and edits RGBA layers for transparent compositing, and is compatible with both diffusers and ComfyUI. SGLang-Diffusion added same-day support for the model with no quantization required; on a single RTX 4090 24GB, 1024x1024 image generation takes 18.7 seconds and image editing takes 21.7 seconds.
Qwen3.8-LiveTranslate, a real-time simultaneous interpretation model, was released supporting 60 languages. Built on an Interleave architecture, it reduces average lagging from 2.8 seconds to 2.3 seconds. The model adds real-time speaker diarization that distinguishes speakers in multi-party speech and preserves each speaker's voice via stable voice cloning, a synchronized bilingual display, and long-context disambiguation leveraging conversation history.
How it developed
Apsara Conference 2026: Alibaba CEO Eddie Wu announced 1,000x cognitive capacity goal, Zhenwu V900 chip specs, and 20GW data-center target
SGLang-Diffusion added same-day support for Qwen-Image-2.1; Hugging Face Spaces demo live
Qwen-Image-2.1 open-weight 7B model released for image generation and editing
Qwen3.8-LiveTranslate released, supporting 60 languages with reduced latency and speaker diarization
Alibaba reported deploying over 100 engineers to embed Qwen as a task-executing agent in Taobao and other retail apps
Sources
- Introducing Qwen Intelligence, bringing personal intelligence within everyone's reach. 📱✨
- ⚡ Meet Qwen-Audio-3.1! ASR, TTS & Realtime are fully upgraded, joined by two new models: TTS-Next for audio creation and…
- Thanks @sgl_project for the day-0 support! 🙌 SGLang-Diffusion now serves Qwen-Image-2.1: text-to-image generation, multi…
- Qwen-Image-2.1 × @HuggingApps: live demo on Spaces! 🖼 One single checkpoint for generation and editing. Try it in your b…
- Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weig…
- Meet Qwen3.8-LiveTranslate, Qwen's next-generation real-time simultaneous interpretation model! 📢
Want this in your inbox?
I send one email each morning with the stories that moved. If you would rather just read here, that works too.
Subscribe free