Google has released two variants of its Gemini 3.8 model, splitting the product line into a fast, conversational version and a slower one built for harder problems. The move formalizes a divide that's become standard across the AI industry: one mode for talking, another for thinking.

Gemini Live is the real-time interaction layer Google has been building out since Gemini 2.0, designed for voice and video conversations where speed matters more than depth. You talk, it responds almost instantly, and the tradeoff is that it isn't built to grind through multi-step logic or catch subtle errors in a complex task. The 3.8 update appears to extend that line with incremental improvements to responsiveness and multimodal handling โ€” the kind of update that rarely gets a headline of its own.

The more notable piece is Extended Thinking, a mode that deliberately slows the model down to work through problems in stages before answering. This isn't a new concept. OpenAI introduced reasoning-focused models with the o1 series in late 2024, and Anthropic added an extended thinking mode to Claude in 2025. Google's version brings Gemini into the same pattern: pay a latency cost, get a more reliable answer on tasks involving math, code, planning, or anything with several dependent steps.

What's new here isn't the idea of a reasoning mode โ€” it's Google folding that capability directly into its Live, real-time product line rather than keeping it as a separate, standalone model. That's a meaningful product decision. It suggests Google wants users to be able to toggle between fast and careful thinking within the same conversational tool, rather than switching to an entirely different app for harder tasks.