Qwen Audio Agent
Agent Presence
Real conversation should not leave you waiting after a single sentence, nor should it grind to a halt just because the Agent is looking something up, calling a tool, or working on a task.
Conversation should keep flowing, and the Agent should always be present.
That is why we built qwen-audio-agent—a realtime voice runtime that keeps Agents talking, working, and present. Whether chatting with you, thinking through a problem, or working on a task, your Agent remains in the conversation. It listens, responds, and when the task is complete, naturally tells you:
"It's ready."
News
-
2026-08-20 · v1.11.0 🧩 Adds embeddable Gateway and Realtime Provider extensions; 🛠️ supports installing and managing Agent Skills; 📎 adds multimodal input to the TUI; 🎨 links pet animations to runtime states.
-
2026-08-13 · v1.10.0 🐋 Added experimental DeepSeek Harness backend support with one-click installation.
-
2026-08-13 · v1.9.0 🧩 Desktop task cards show live Agent progress; 🔎 backend Agent selection is clearer and searchable; 🎙️ supports Qwen3.5-Omni Realtime frontend integration.
-
2026-08-09 · v1.8.0 🆕 Adds Qwen Code backend; 🔧 fixes known issues.
-
2026-08-07 · v1.7.0 🎨 The orb opens up custom skins — import your own look, compatible with pet packs from the Awesome Codex Pet community gallery; 🪟 improved Windows backend Agent startup.
-
2026-08-06 · v1.6.0 🪟 Desktop app now officially supports Windows; 🧠 adds invisible memory with automatic extraction after each session.
-
2026-08-05 · v1.5.0 ⏰ Adds scheduled reminders and progress reporting; 🗣️ adds the voice wake word ("你好千问"); 🐧 desktop build support for Linux; the desktop app now uses a data directory isolated from the CLI.
-
2026-08-04 · v1.4.0 🧠 Adds personalized rules and checklist management; desktop app supports auto-sleep and shortcut wake.
-
2026-08-03 · v1.3.0 🎙️ Adds 🤗 speech-to-speech frontend integration, supporting fully local VAD, STT, LLM, and TTS.
-
2026-08-01 · v1.2.0 ⚡ Desktop app adds auto-update, faster startup, and improved backend Agent detection.
-
2026-07-31 · v1.1.0 🤝 Adds Kimi Code CLI backend with native ACP integration.
-
2026-07-30 · v1.0.0 🚀 First stable release, introducing a macOS desktop app with a built-in Gateway.
-
2026-07-28 · v0.9.0 🌍 Project officially open-sourced; backend Agents unified under the ACP architecture.
Conversation Continues, Tasks Too
Conversation doesn't stop for background tasks; when a task completes, the result naturally returns to the current conversation:
https://github.com/user-attachments/assets/ab570531-8da9-4af4-93fa-244bb6614c05
Core Features
-
Full-duplex realtime voice interaction, natural interruption, and sustained multi-turn conversation
-
DashScope Qwen Audio and Qwen3.5 Omni Realtime model selection from one shared model catalog
-
One-click selection of your preferred coding Agent, reusing existing tools, MCP, and Skills
-
Frontend conversation and background tasks run in parallel; ask about progress or cancel at any time
-
Create multiple independent tasks executed asynchronously by the backend Agent, with continuous status tracking
-
Task results automatically return to the current conversation, supporting follow-up questions and modifications
-
WebUI, terminal TUI, and desktop floating orb (macOS / Windows / Linux)
-
Desktop auto-sleep disconnects cloud Realtime without stopping submitted tasks; wake with a configurable shortcut or the local wake word
-
Per-user long-term personalization and cross-session memory
Architecture

Questions that can be answered directly are answered immediately; when tools or sustained processing are needed, the task is delegated to the backend Agent. Throughout, the user always faces the same assistant.

For the full design and module breakdown, see the architecture document.
Agent Support
Backend Agent Integration Setup Rating
None N/A Frontend-only mode, no config needed ★★★★★
OpenCode Native ACP One-click install + Bailian config ★★★★★
OpenClaw Built-in ACP bridge One-click install + Bailian config ★★★★★
Qoder Native ACP One-click install, user config required ★★★★★
Qwen Code Native ACP One-click install, user config required ★★★★☆
Kimi Code Native ACP One-click install, user config required ★★★★★
Hermes Native ACP One-click install, user config required ★★★★☆
CodeBuddy Native ACP One-click install, user config required ★★★★☆
Codex External ACP adapter One-click install (base + adapter), user config required ★★★★☆
Claude Code External ACP adapter One-click install (base + adapter), user config required ★★★★☆
DeepSeek Native ACP One-click install, DeepSeek API key required ★★★★☆
Pi External ACP adapter One-click install (base + adapter), user config required ★★★★☆
Ratings reflect current integration completeness, compatibility, and verification level: five stars indicate a thoroughly tested recommended integration; four stars indicate active development or not yet fully verified. For detailed configuration and capability boundaries, see the configuration guide.
Installation
Requires Node.js 22.22.2+ or 24.15.0+, npm 10+. One-click install (recommended):
npm install -g qwen-audio-agent
For building from source, installing from GitHub, and obtaining a DashScope API Key, see the installation guide.
Quick Start
- Create your config and fill in the API Key:
qwenaudio config
DASHSCOPE_API_KEY=your-key
# Voice frontend model: Omni Flash/Plus or Audio Flash/Plus (Audio Plus is default)
QWEN_AUDIO_REALTIME_MODEL=qwen-audio-3.0-realtime-plus
# Backend Agent: optional, leave empty or set to none for frontend-only mode
AGENT_PROTOCOL=openclaw
# Backend model: can be empty; if empty, uses the Agent's own user config
QWEN_AUDIO_AGENT_BACKEND_MODEL=qwen3.7-max
Uses DashScope realtime voice frontend by default; alternatively, switch to a local speech-to-speech frontend, no cloud API Key needed.
qwen3.5-omni-flash-realtime and qwen3.5-omni-plus-realtime
accept text, audio, and image at the model level. This release transports text and
audio only; image/frame and native-video transport remain disabled until their client and
Gateway paths are implemented.
The Desktop app or qwenaudio config set --realtime-model <id> configures the single
Gateway-wide model. Restart the Gateway after a CLI change. WebUI and TUI display the active
model but do not override it.
- Start the Gateway, then open another terminal to start the TUI (or use
qwenaudio webuifor the browser UI):
qwenaudio # Terminal 1: Gateway
qwenaudio tui # Terminal 2: TUI
For full configuration options, speech-to-speech frontend setup, and TUI platform notes, see quick start, voice frontends, and TUI notes.
Examples
This repository includes a smart cockpit voice Agent example with vehicle control, navigation, music, weather, web search, flash-buy workflows, and a car UI:
cp examples/car/.env.example examples/car/.env.local
npm install --prefix examples/car/server
npm install --prefix examples/car/react-app
npm run example:car:server # Terminal 1: car Agent server
npm run example:car:web # Terminal 2: car UI
See examples/car for details.
Desktop App
The desktop app provides a floating voice orb that stays on your desktop, with a built-in Gateway, automatic idle sleep, a configurable wake shortcut, and a local voice wake word. Sleep disconnects the Realtime frontend while keeping the app, Gateway, backend Agent, and submitted tasks alive; it is not the same as quitting or restarting the deskto