Alibaba has released Qwen Audio 3.0 Realtime, with the Plus variant debuting as the new #1 model on the Artificial Analysis Speech to Speech Index at 84.1%, ahead of GPT-Realtime-2.1 High at 79.1%
Released earlier this month, Qwen Audio 3.0 Realtime is @Alibaba_Qwen's flagship native Speech to Speech model, available in two variants: Plus and Flash. Qwen Audio 3.0 Realtime Plus leads on all three component benchmarks comprising the Artificial Analysis Speech to Speech Index, Big Bench Audio for Speech Reasoning, Full Duplex Bench for Conversational Dynamics, and Tau Voice for Agentic Performance. We tested the China-hosted endpoints on Aliyun (Alibaba Cloud).
Key takeaways: ➤ Speech to Speech Index: Qwen Audio 3.0 Realtime Plus is the new leader at 84.1%, ahead of GPT-Realtime-2.1 High (79.1%) and GPT-Realtime-2 High (77.2%). The Flash variant comes in at 4th at 76.3%. ➤ Speech to Speech Index by Benchmark: On Big Bench Audio, the Plus variant achieves 99.2%, up ~0.5 percentage points from the previous best of 98.7% (Alibaba Qwen3.5 Omni Plus Realtime), with Flash variant scoring 96.1%. On Tau Voice, the Plus variant currently leads with a score of 54.6%, ahead of Grok Voice Think Fast 1.0 at 52.1%. On Full Duplex Bench, the Plus variant leads our Full Duplex Bench subset at 98.4%, with Flash at 96.9%, both ahead of the best non-Alibaba model, GPT-Realtime-2 (Minimal) at 96.1%. ➤ Speed: The Plus variant records an average Time to First Audio of 4.02 seconds on Big Bench Audio, with Flash at 4.16 seconds, among the slowest models on our leaderboard, and well behind GPT-Realtime-2 (Minimal) at 1.10 seconds and GPT-Realtime-2 (High) at 1.14 seconds ➤ Price: Plus costs $4.42 per hour of input audio on our Big Bench Audio subset, more expensive than GPT-Realtime-2 High ($4.14) and ~2.4x cheaper than GPT-Realtime-2.1 High ($10.75). The average cost for Flash variant is $4.77, higher than the Plus variant despite lower list prices, driven by comparatively more verbose responses.