What OpenAI's Multilingual Retail Voice Case Shows — Operating Metrics, Not a Demo
30,000 users in two weeks, 92% positive survey responses. What is new in the avatarin / Yamada Denki case is the kind of metric disclosed, not its size.
Thirty thousand people used it in two weeks.
That figure comes from the avatarin case study OpenAI published on its own blog. avatarin ran a GPT-Realtime voice agent for shoppers of Japanese electronics retailer Yamada Denki, 24/7 and multilingual, reporting 30,000 users over two weeks with 92% of survey responses positive.
1. Two Numbers, With a Defined Character
- 30,000 users over two weeks — scale and period stated together
- 92% positive survey responses — respondent count and questions not disclosed
Both were published by OpenAI as its own customer story, not as an independent industry benchmark. If they travel into an adoption review, that attribution travels with them. A satisfaction figure with an unknown response rate makes a poor baseline.
2. What Changed Is the Kind of Metric, Not Its Size
Voice AI announcements generally lead with conversational quality — how natural the response sounds, whether the agent recovers when a caller talks over it. This one leads with user count, operating period, and running 24/7 across languages. The criterion is moving from "does it sound right" to "how many used it."
A demo shows whether one call goes well. Operating metrics show whether the thirty-thousandth call matched the first.
3. Always-On Was Framed as a Cost Design Problem
OpenAI also stated that it worked with avatarin to structure complex prompts and to optimize API costs for an always-on voice service. That puts continuous multilingual operation on record, from the vendor side, as a cost design problem before a technical one. Realtime models bill against talk time, so idle handling and session policy decide unit cost.
4. Reading This From Korea
- The demand exists here too. Tourist-district stores and cross-border e-commerce already take foreign-language inquiries after hours
- Set adoption criteria on operating metrics. Rather than how natural the demo felt, verify three lines at quote stage: language coverage, response time, unit cost under continuous operation
- Re-measure satisfaction yourself. 92% is someone else's survey; a comparison exists only if you set the response rate and questions