An AI recording studio for creators

Qwen-Audio-3.0-TTS Online Text to Speech & Voice Cloning

Generate natural, expressive speech online with Alibaba Qwen-Audio-3.0-TTS. Create multilingual text-to-speech audio, clone voices, use Chinese dialects, emotion tags, and instruction control for videos, podcasts, games, audiobooks, and product narration.

Trusted by 12,000+ creators

Voice workbench

Turn your words into a voice

Shape the script, voice, and delivery in one calm, focused workspace.

0/ 2000
Model

More characters per generation·Faster generation·More voice slots

Upgrade

Model guide

How do Qwen-Audio-3.0-TTS, Qwen3-TTS, and CosyVoice differ?

All three are speech synthesis model families available through Alibaba Cloud Model Studio, but they are separate families rather than consecutive versions of one model. Their voice sources, voice cloning, voice design, instruction control, and integration methods differ.

Qwen3-TTS

Qwen3-TTS belongs to the Qwen-TTS family and includes separate models for standard synthesis, instruction control, voice design, and voice cloning. It uses different model names and some different integration methods from Qwen-Audio-3.0-TTS, so voices and API parameters are not interchangeable.

CosyVoice

CosyVoice is another independent speech synthesis family with real-time and non-real-time generation. Different versions provide voice cloning, voice design, and instruction control, with integration options designed for real-time interaction, voice clients, and more demanding network conditions.

DS Speech currently focuses on the Qwen-Audio-3.0-TTS online experience, so you can select voices, clone a voice, control delivery, and generate audio without configuring models, APIs, or storage yourself.

Online text to speechOnline voice cloning

Simple, transparent pricing

Pick a plan that fits your rhythm

Membership allowance and credit packs reset each calendar month. Every plan uses USD pricing.

The first-purchase offer can be used once per user. Yearly plans show the full annual charge; the monthly equivalent is for comparison only.

Creator reviews

Creators across industries use it to produce voice content

From short videos and podcasts to games, courses, and branded content, creators use DS Speech to produce natural, expressive AI narration faster.

Qwen-Audio-3.0-TTS is fast enough for daily production. I can revise a script and hear a stable voice result in seconds instead of booking another recording session.

林澈林澈Short-form Video Director

The emotional delivery does not feel robotic. A short instruction gives podcast intros, ads, and transitions a much more intentional tone.

Mika ChenMika ChenPodcast Producer

The public voice library covers character narration, NPC dialogue, and tutorial prompts, so reviews no longer rely on temporary machine voices.

周远周远Game Narrative Designer

Dialect and multilingual text to speech make regional training content sound closer to the audience and less generic than ordinary TTS.

Nina HayesNina HayesTraining Producer

We use voice cloning for course series. When an instructor cannot record again, updates still keep a consistent voice across lessons.

何嘉宁何嘉宁Corporate Learning Producer

Qwen-Audio-3.0-TTS is fast enough for daily production. I can revise a script and hear a stable voice result in seconds instead of booking another recording session.

林澈林澈Short-form Video Director

The emotional delivery does not feel robotic. A short instruction gives podcast intros, ads, and transitions a much more intentional tone.

Mika ChenMika ChenPodcast Producer

The public voice library covers character narration, NPC dialogue, and tutorial prompts, so reviews no longer rely on temporary machine voices.

周远周远Game Narrative Designer

Dialect and multilingual text to speech make regional training content sound closer to the audience and less generic than ordinary TTS.

Nina HayesNina HayesTraining Producer

We use voice cloning for course series. When an instructor cannot record again, updates still keep a consistent voice across lessons.

何嘉宁何嘉宁Corporate Learning Producer

I use AI voice cloning to test ad reads quickly and compare several scripts and delivery directions in one afternoon.

Alex MorganAlex MorganGrowth Lead

Long-form text generation is easy to manage, and emotion tags help shape each chapter before the final pacing pass.

宋怡宋怡Audiobook Editor

A cloned character voice stays consistent across episodes, which makes ongoing series much easier to produce.

Kenta MoriKenta MoriAnimation Channel Producer

We use online text to speech for product explainers and launch teasers, even when the final script changes close to release.

陈玥陈玥Brand Content Strategist

The voice plaza has enough range for explainers, shorts, and product clips without starting a voice search from scratch every time.

Oliver GrantOliver GrantVideo Producer

I use AI voice cloning to test ad reads quickly and compare several scripts and delivery directions in one afternoon.

Alex MorganAlex MorganGrowth Lead

Long-form text generation is easy to manage, and emotion tags help shape each chapter before the final pacing pass.

宋怡宋怡Audiobook Editor

A cloned character voice stays consistent across episodes, which makes ongoing series much easier to produce.

Kenta MoriKenta MoriAnimation Channel Producer

We use online text to speech for product explainers and launch teasers, even when the final script changes close to release.

陈玥陈玥Brand Content Strategist

The voice plaza has enough range for explainers, shorts, and product clips without starting a voice search from scratch every time.

Oliver GrantOliver GrantVideo Producer

Longer scripts fit into each generation, so the text needs fewer cuts and the finished delivery sounds more connected.

梁书瑶梁书瑶Educational Creator

Multilingual AI narration is much faster now. We can compare English, Chinese, Japanese, and Korean versions in the same review.

Sophie CarterSophie CarterBrand Strategist

Regional accents matter for local content. Comparing standard Mandarin and dialect versions helps the team choose what feels closest to the audience.

高铭高铭Local Content Manager

Emotion tags make it quick to separate tense dialogue from calm guidance, giving each character more expressive range.

Haruto SatoHaruto SatoGame Audio Director

Course updates depend on speed. I can revise a lesson and regenerate clean narration without scheduling a full recording session.

Rachel MooreRachel MooreCourse Creator

Longer scripts fit into each generation, so the text needs fewer cuts and the finished delivery sounds more connected.

梁书瑶梁书瑶Educational Creator

Multilingual AI narration is much faster now. We can compare English, Chinese, Japanese, and Korean versions in the same review.

Sophie CarterSophie CarterBrand Strategist

Regional accents matter for local content. Comparing standard Mandarin and dialect versions helps the team choose what feels closest to the audience.

高铭高铭Local Content Manager

Emotion tags make it quick to separate tense dialogue from calm guidance, giving each character more expressive range.

Haruto SatoHaruto SatoGame Audio Director

Course updates depend on speed. I can revise a lesson and regenerate clean narration without scheduling a full recording session.

Rachel MooreRachel MooreCourse Creator

Frequently Asked Questions

1

What is Qwen-Audio-3.0-TTS, and how good is its voice cloning?

Qwen-Audio-3.0-TTS is a high-performance speech synthesis family available through Alibaba Cloud Model Studio. It supports online text to speech, voice cloning, voice design, and instruction-based delivery control. It is a separate model family from Qwen3-TTS and CosyVoice. DS Speech focuses on making Qwen-Audio-3.0-TTS voice cloning and AI speech generation easy to use online.

2

How do I start using AI voice cloning?

Enter text in the workbench, choose a preset voice, or upload a clear reference audio clip to save your own voice model, then generate a preview.

3

Can generated voices be used commercially?

Yes. Before using someone else's voice, make sure you have legal authorization from the voice rights holder. If you use your own voice, or a voice generated by you that does not involve third-party rights, it can be used for commercial projects without additional commercial-use restrictions from us.

4

Which languages are currently supported?

Chinese (Mandarin, Guangdong dialect, Chongqing dialect, Northeastern dialect, Gansu dialect, Guizhou dialect, Zhejiang dialect, Hebei dialect, Henan dialect, Hubei dialect, Hunan dialect, Jiangxi dialect, Ningbo dialect, Ningxia dialect, Qingdao dialect, Shaanxi dialect, Shanxi dialect, Shandong dialect, Shanghai dialect, Sichuan dialect, and Yunnan dialect), English, Japanese, Korean, Russian, French, German, Portuguese, Thai, Indonesian, Vietnamese, Spanish, Italian, Malaysian, Filipino, and Arabic.

5

What are the requirements for reference audio?

Supported formats: WAV, MP3, and M4A. Recommended duration is 10~20 seconds, with a maximum of 60 seconds. File size must be ≤ 10 MB, and sample rate must be ≥ 16 kHz.

For stereo audio, only the first channel is processed, so make sure the first channel contains valid speech.

The audio must contain at least 5 seconds of continuous, clear reading without background sound. Other parts may only contain short pauses of ≤ 2 seconds. Avoid background music, environmental noise, or other voices. Use normal spoken audio; do not upload songs or singing recordings.

6

What usage restrictions apply?

Do not use the product for fraud, impersonation, harassment, hate, obscene or pornographic content, or any illegal scenario. If discovered, your account rights will be terminated immediately.

7

How do membership allowances and credit packs reset?

Membership allowance and separately purchased credit packs are cleared each calendar month. Please plan usage within the current month.

8

Can I control emotion, dialect, or character style?

Yes. Tags and instruction prompts can control some emotion, dialect, and character styles.

Instruction examples:

Standard broadcast style: clear and precise articulation.

Young energetic female voice: faster pace with an upward intonation, suitable for fashion products.

Calm middle-aged male: slow pace and deep magnetic tone, suitable for news or documentary narration.

Chinese dialect instruction example:

请用河南话表达

Tag example:

[excited]The weather is great today![laughing]Let's go out together!

9

Will my voice models be cleared after my membership expires?

No. After your membership expires or you downgrade from Pro to Plus, your created voice models will be retained and not deleted. However, you can only use the most recently created voice models within the quantity supported by your current membership level.

For example, if you were a Pro member and created 30 voice models, after downgrading to Plus you can continue using the 10 most recently created voice models. The remaining models are retained but temporarily unavailable.

When you reactivate Pro, all 30 original voice models will become available again.

10

How long are text-to-speech audio files retained?

Audio files generated with text-to-speech are retained for only 3 days. After that, they can no longer be played or downloaded, so please download and store them promptly.