본문으로 건너뛰기
BringTalk
프로그램▼
워크샵2일 — 직접 만들어보며 감을 잡습니다POC프로젝트★ 대부분 여기서 시작6주 — 빠르게 데이터로 검증합니다AX프로젝트검증한 것을 전사의 시스템으로
콘솔Alpha
인더스트리▼
자동차광고 리드, 시승, 견적, 정비 예약까지 콜을 매출 흐름으로 연결합니다.인테리어·가구상담 예약, 견적 일관성, 설치 후 CS를 한 흐름으로 묶습니다.보험 GA리드 후속, 보장 안내, 갱신 콜을 컴플라이언스 기준으로 운영합니다.금융·캐피탈납부 일정, 신청 상태, 동의 확인 콜을 신뢰 기반으로 처리합니다.통신·인터넷해지 방어, 장애 접수, 요금제 상담을 반복 가능한 운영으로 만듭니다.여행·항공·호텔변경, 지연, 리워드 문의를 고객 맥락에 맞춰 이어갑니다.법무·회계·세무첫 상담 접수와 일정·서류 안내로 전문가 시간을 회수합니다.의료·치과예약, 시술 상담, 진료 외 시간 응대를 놓치지 않고 받습니다.부동산 중개매물 매칭과 임장 일정 조율로 거래 기회를 지킵니다.교육·유학·학원상담 신청 후속과 등록 전환을 부모·학생 양쪽 톤으로 응대합니다.물류·렌탈·모빌리티배송 추적, 렌탈 일정, 차량 상태 조회 콜을 자동화합니다.B2G·공공민원 1차 분류와 인증·정책 안내로 상담 인력을 비웁니다.
솔루션▼
Vapi공식 파트너Global Top 3 Voice AI Platform
블로그
상담 신청
프로그램
워크샵POC프로젝트AX프로젝트
콘솔Alpha
인더스트리
자동차인테리어·가구보험 GA금융·캐피탈통신·인터넷여행·항공·호텔법무·회계·세무의료·치과부동산 중개교육·유학·학원물류·렌탈·모빌리티B2G·공공
솔루션
Vapi
블로그
상담 신청
설계

How OpenAI Cut Voice Latency by Removing Steps

In OpenAI's August 2026 account of building its realtime voice system, the two biggest numbers sat outside the model: connection round trips fell from six to one, and the turn detector was removed from the audio path.

Moon Kim·August 10, 2026·2분 읽기

목차

  1. The Largest Number Lands Before Anyone Speaks
  2. The Next Move Was Deleting a Component
  3. The Third Change Was a Rewrite
  4. Reading This From a Contact Centre
  5. When This Does Not Transfer

When a call feels sluggish, the model is the first thing anyone looks at. In OpenAI's August 2026 account of building its realtime voice system, the biggest numbers from six months of work sit outside the model.

All figures below are OpenAI's own, against its own previous system. Not an independent benchmark.

The Largest Number Lands Before Anyone Speaks

OpenAI built a transport protocol called WARP that cut the network round trips needed to start sending audio and data from six down to one. With Instant Connect on top, a client can open a session with one UDP packet.

Dashboards rarely capture this stretch. Most teams measure question-to-answer time, and those six round trips pass before that clock starts.

The Next Move Was Deleting a Component

Calling GPT-Live its third-generation voice system, OpenAI wrote that it removed the turn detector from the audio path. The piece deciding whether the other person had stopped talking is gone.

Once the model listens and speaks at once, there is no decision left to make. Nothing was made faster. A step was deleted.

What OpenAI removed at each of three layers: connection, pipeline, implementation
What OpenAI removed at each of three layers: connection, pipeline, implementation

None is a swap for a faster part. OpenAI's own disclosures.

The Third Change Was a Rewrite

The media frontend and inference logic moved from Python asyncio to Go, reported this way.

new system p95  =  previous system p50

No absolute millisecond figures were released, only percentiles. The slow tail moved to where the middle used to sit.

Tool calls and heavier reasoning also moved onto a separate asynchronous path, so a slow backend cannot stall the conversation.

Reading This From a Contact Centre

None of the three answers which component to swap for a faster one. That changes where to look.

  1. Check whether connection setup is measured on its own. A dashboard tracking only response time has no column for it
  2. Count the steps in the path. Removing one beats optimising it
  3. Check whether external lookups hold the call. If speech stalls while a CRM query runs, the paths were never separated

There are two ways to reduce latency. Do it faster, or do not do it.

When This Does Not Transfer

Two of these travel badly.

  • Dropping the turn detector only works when the model listens and speaks at once. A staged STT → LLM → TTS stack has nothing else to tell it a sentence ended. Remove endpointing there and it cuts people off
  • WARP is a WebRTC protocol, as its name says. For calls arriving over the public telephone network, the carrier leg is not yours to change, and six-to-one does not carry over
Four things to check before adopting any of this
Four things to check before adopting any of this

Which items apply depends on how the stack is built.

At BringTalk we count a path's steps before reading its response time. While a step can still be removed, changing the model comes after.

Baseline: In OpenAI's own account, the largest numbers were connection round trips going from six to one, and the turn detector's removal. None of the three changes was a faster part.

이 글 공유하기
XLinkedIn

READ NEXT

함께 보면 좋은 글

설계

Voice AI Retry Design — retryOnFail, Idempotency Keys, and onError

September 3, 2026
설계

Where Phone Booking Automation Ends

September 2, 2026
설계

Traditional IVR, Visual IVR, Digital ARS, Voice AI — Four Systems With Different Branch Logic

September 1, 2026
우리 콜에서는?

같은 전환을 한국어 콜 운영에서 — 6주 안에 숫자로 확인하세요.

6주 POC 상담Vapi 도입 상담
BringTalk

콜 운영에 들어가 음성 AI 에이전트를 6주 만에 실험 가능한 시스템으로 구축합니다.

탐색
  • 브링톡 콘솔 Alpha
  • 인더스트리
  • Vapi 파트너십
  • 블로그
프로그램
  • 워크샵
  • POC프로젝트
  • AX프로젝트
연락
  • contact@bringtalk.ai
  • 070-5275-3800
  • 상담 신청
개인정보처리방침이용약관개인정보 문의
© 2026 BringTalk · Voice Agent StudioEvery call becomes revenue.