Cryptelio

Markets

Google Launches Gemini 3.8 Text-to-Speech Models for Enhanced Voice Applications

Cryptelio Editorial Published 23 Sep 2026 · 15:45 UTC

Google has introduced its new Gemini 3.8 text-to-speech models, enhancing the capabilities of creators and developers to produce expressive, natural, and multilingual voices on a large scale. This release is part of the broader Gemini audio family, which includes models like Gemini 3.8 Live and Gemini 3.8 Flash, aimed at real-time voice interactions and coding tasks, respectively.

The introduction of Gemini 3.8 text-to-speech models appears to enhance Google’s competitive position in AI technology development. This release is part of a broader strategy to expand the Gemini audio stack, suggesting a comprehensive approach to AI-driven audio applications. Market participants interpret this development as indicative of Google’s ongoing advancements in AI, potentially influencing perceptions of the company’s AI model performance.

Key Features of Gemini 3.8

  • Enhanced voice customization with three tiers of voice selection.
  • Supports up to 130 languages, catering to global markets.
  • Includes both single-speaker and multi-speaker audio generation.
  • Features a custom voice design system allowing natural language prompts for voice creation.

The TTS launch follows the release of Gemini 3.8 Live models, designed for real-time speech interactions, showcasing Google's commitment to advancing AI-driven voice technologies. Market participants will likely monitor how these models perform relative to competitors, such as those from Anthropic and OpenAI, as well as their impact on the Chatbot Arena LLM Leaderboard.

Latest Updates on ChatGPT Voice

  • ChatGPT Voice has been updated to include plugin support and three new GPT-6 models: Astra, Sol, and Luna.
  • Astra is OpenAI's most capable model, designed for coding and complex reasoning, while Sol and Luna are optimized for everyday use cases.
  • As of September 10, 2026, Pro plan users can select between GPT-5.6 Sol or GPT-6 Astra for voice interactions.
  • GPT-6 Sol and GPT-6 Luna launched on September 22, 2026, with API pricing significantly reduced—Sol at $2 per million input tokens and $10 per million output tokens, and Luna at $0.10/$0.50.
  • OpenAI has adjusted daily usage caps for GPT-Live, allowing Plus subscribers up to 3 hours of voice interaction and Pro users 15 hours per day.
  • Luna is available to Free and Go tier users through the desktop app, which gained GPT-Live support in July 2026.
  • The new pricing strategy for Sol and Luna aims to reshape developer behavior by offering more cost-effective solutions for various tasks.

New Developments in Google Gemini 3.8 Text-to-Speech Models

  • Gemini 3.8 Flash TTS achieved a score of 89.5% on the Artificial Analysis Pronunciation Robustness Benchmark, surpassing its predecessor, Gemini 3.1 Flash TTS, which scored 88.2%.
  • The model also ranked first on the Hume AI Voice Design Benchmark with an overall score of 71.4, and specifically scored 60.8 for accent reproduction.
  • Gemini 3.8 Flash TTS secured both the first and second positions on the Hume AI Overall Quality Index.
  • Launched on September 23, Gemini 3.8 Flash TTS includes a lighter variant called Flash-Lite TTS, both available through the Gemini API and Google AI Studio.
  • The model supports generative voice design from text prompts and can replicate a specific voice from approximately 30 seconds of audio.
  • It is capable of handling multi-speaker interactions, allowing a single model to voice entire conversations between different characters.
  • Language support exceeds 100 languages, including various regional accents, with studio-grade voice fidelity and fine-grained control through metadata and inline tags.

FAQ

What are the key features of the Gemini 3.8 text-to-speech models?

The key features include enhanced voice customization with three tiers of voice selection, support for up to 130 languages, single-speaker and multi-speaker audio generation, and a custom voice design system that allows for natural language prompts for voice creation.

How does Gemini 3.8 compare to previous models?

Gemini 3.8 builds on previous models by offering improved expressiveness, naturalness, and multilingual capabilities, making it more suitable for a wider range of applications in voice technology.

What is the significance of the Gemini audio family?

The Gemini audio family, including models like Gemini 3.8 Live and Gemini 3.8 Flash, represents Google's comprehensive approach to AI-driven audio applications, enhancing real-time voice interactions and coding tasks.

How many languages does Gemini 3.8 support?

Gemini 3.8 supports up to 130 languages, catering to a diverse global market.

What impact might Gemini 3.8 have on the AI market?

The launch of Gemini 3.8 is expected to enhance Google's competitive position in AI technology development, potentially influencing perceptions of the company's AI model performance and impacting the competitive landscape against companies like Anthropic and OpenAI.

Read story →