Google is teaching AI to understand language, not just translate text
Google has expanded its language models to cover 300+ languages, combining translation, speech recognition and emotion detection into a single Gemini-based system.
Google has shared how its technologies now work with more than 300 languages, spoken by over 7 billion people — 86% of the world's population. The company calls this an interim milestone in years of work, from the classic Google Translate to the new Gemini family of models, which understand not just words but intonation, pauses, language-mixing within a single phrase, and the speaker's emotion.
For a business owner, this isn't an abstract piece of linguistics news. Speech and text AI is built into search, maps, advertising tools, voice assistants and translators that your customers — and you — use every day. If the technology has become better at understanding natural speech across dozens of languages, that changes how people search for products by voice, how automatic translation of reviews and descriptions works, and how ads sound in different countries. Some of this brings practical benefits; some of it is worth examining closely before celebrating.
What's changed: from translating text to understanding speech
As Google explains, speech recognition used to follow a rigid pipeline: audio was converted into text, the text was processed, and then it was converted back into speech. That pipeline worked, but it lost the most important things — tone, pace, emotion, context. People don't speak in perfectly grammatical sentences: they laugh, interrupt each other, stumble over words, and mix languages within a single phrase — as in Spanglish or Hinglish.
To account for this, the company moved to what it calls “native audio intelligence”: models, including Gemini, trained to process sound directly, without converting it into flat text first. According to Google, this has produced concrete tools:
Gemini 3.5 Live Translate — real-time speech translation across 70 languages and more than 2,000 language pairs, recognizing language switching and emotional nuance.
Gemini 3.5 Transcribe — the company's most accurate speech-to-text model to date: it turns raw audio into clean text even in noisy environments or with complex terminology. It powers the Rambler feature in Gboard on Android, which removes filler words, fixes grammar and punctuation, lets you edit text by voice, and lets you switch between languages on the fly.
For businesses, this means voice search, voice notes, and automatic transcription of calls and meetings are becoming noticeably more accurate. This is especially useful when customers speak with an accent, mix languages (say, Russian with Kazakh or Ukrainian), or talk in a noisy environment, which is common in retail and services.
The “1,000 Languages” initiative, and why this isn't just about rare dialects
Google has announced a goal of supporting the 1,000 most widely spoken languages in the world, not just the ones where AI already performs well. The problem is that the internet and most training data skew heavily toward a handful of dominant languages, while the rest are poorly represented or entirely absent from the digital world.
To tackle this, Google trained its Universal Speech Model on 12 million hours of audio and applied cross-lingual transfer learning: the model learns from languages with abundant data and applies the patterns it finds to languages with little data. According to the company, this work builds on 25 years of open research and more than 400 scientific publications on speech technology.
At the same time, Google is building partnerships with local communities to collect data “on the ground”:
WAXAL — an open speech dataset covering 27 languages of Sub-Saharan Africa, spoken by more than 100 million people across 26+ countries. The project was built with Makerere University and Digital Umuganda and specifically captures tonal features and the rhythm of conversational speech, which are usually lost in standard datasets.
Project Vaani — together with the Indian Institute of Science (IISc) and the Bhashini organization, the project has collected more than 30,000 hours of speech across 109 languages from over 155,000 speakers; the data is mapped not by language but by region across India.
Amplify Initiative — more than 1,600 local experts and 20 universities across four continents, including UFMG in Brazil, IIT Kharagpur in India, and Makerere University in Uganda, have collected 15,000 multimodal data points reflecting local characteristics.
Another tool is Language Explorer, a visualization of the LinguaMeta database. Google calls it the largest open repository of language data in the world, covering more than 7,000 spoken, written and signed languages.
None of this directly involves languages of the CIS, but the underlying principle matters: in places where AI used to perform poorly due to lack of data, quality will improve faster thanks to cross-lingual transfer.
Working without internet access, and on basic phones
Another part of the announcement concerns accessibility. According to Google, citing Forbes research, more than 3 billion people still lack reliable internet access. Technology only truly works where it's accessible in real-world conditions — not just in demos running on fast Wi-Fi.
To address this, Google released TranslateGemma — a family of lightweight translation models built on Gemini and trained on 55 languages. The model runs directly on the device, without the cloud and without an internet connection, which matters in regions with poor network coverage.
Even lightweight models require a certain level of device capability, and hundreds of millions of people still use basic feature phones. Here, Google supports Viamo and its “Ask Viamo Anything” (AVA) service — a Gemini-based voice AI assistant that works over ordinary calls to a feature phone. According to the company, in Rwanda the service has already answered more than 2 million user questions.
Sign language and cultural context
Google also notes that language isn't just dialects and vocabulary — it includes other forms of communication that traditional speech tools typically ignore, forcing people to adapt to the technology rather than the other way around.
One response to this is Sign Language-to-Text (SL2T), a model trained on more than 50 sign languages. It's built into the sign-writing feature in Gboard and into Live Transcribe on Pixel 11, currently for the American Sign Language–English pair. Google describes this as a first step toward making its products more accessible to the 70 million people worldwide who communicate using sign language.
A similar approach has been applied to place-name pronunciation. In New Zealand, the company worked with Māori language experts so that Google Maps pronounces local place names correctly — mispronunciation doesn't just confuse users, it also devalues a place's cultural heritage. Correct, culturally accurate pronunciations are now built directly into the speech synthesis models.
What this means for small businesses in Russia and the CIS
The announcement doesn't mention any products specifically for the Russian language, but the general logic applies to any business whose customers use voice search, automatic translation, or communicate with the company not in formal language but the way people actually speak — with accents, dialect features, and mixed languages.
Things worth doing this week:
Check how your business sounds in voice search. Ask Google Assistant or use voice input to say what a customer might ask: “where to buy [product] in [city],” “how much does [service] cost.” If the answer is inaccurate or your business doesn't come up, that's a signal to update your website descriptions and business listings with simple, conversational wording rather than just formal service names.
If you run an online store or a site with descriptions in several languages (Russian, Kazakh, Ukrainian, Uzbek, and so on), check the automatic translations of reviews and product listings. Recognition quality for mixed speech is improving, and it's worth factoring in when updating content for different audiences.
If your business handles customers over the phone — a call center, support desk, phone orders — evaluate automatic call transcription tools. Recognition accuracy in noisy environments and with specialized terminology (medical, technical, industry-specific) has improved significantly, and this could speed up request processing and call quality control without extra spending on manual transcription.
If your audience is in regions with unstable internet access or people accustomed to basic phones, don't rely on complex voice scenarios as your only channel. Keep text and offline alternatives available: SMS, simple IVR menus, printed materials. Technology for these scenarios is still catching up.
The name of your company, product, or street where your store is located may be mispronounced by navigation apps and voice assistants. This isn't critical for large businesses, but for local retail and services it's worth checking how your address sounds in Google Maps and correcting the transliteration if it distorts the name beyond recognition.
The catch is that AI accuracy gains for “major” languages are happening faster than for languages spoken in Russia and the CIS outside of Russian — Ukrainian, Kazakh, Uzbek, and others. There's objectively less data for these languages, and while Google describes cross-lingual transfer as a way to close the gap, results for a specific region may arrive later than for languages with more speakers and more data.
Bottom line
Google describes a shift from translating text to understanding live speech — with emotion, pauses, language-mixing and cultural context. This shows up in concrete products: Gemini 3.5 Live Translate, Gemini 3.5 Transcribe, TranslateGemma, Sign Language-to-Text, as well as in large-scale data-collection partnerships for languages that used to be poorly represented in the digital world.
For businesses in Russia and the CIS, there's no direct benefit from these specific projects yet — they target other regions. Still, the broader trend is worth factoring into planning for voice search, content auto-translation, and call handling: recognition accuracy for conversational, mixed and non-standard speech is improving, and it's worth checking in advance how your particular business shows up in these systems.