Our approach
Why Audentic builds around OpenAI Realtime speech-to-speech
Our case for a business platform built around OpenAI's native end-to-end speech-to-speech architecture, with knowledge, website deployment, and operator control.
By Audentic · Published
Sources checked . Written by the Audentic team.
A business owner should be able to put their knowledge to work without becoming a voice infrastructure engineer. That is the idea behind Audentic: build a useful voice agent, welcome people on your website, and keep the resulting conversations and requests close to the people running the business.
Our core focus is OpenAI Realtime's native end-to-end speech-to-speech model. It receives audio and produces spoken responses directly. We are building Audentic around that architecture and the business work that makes it useful: knowledge, website deployment, controlled capabilities, conversation review, and customer requests.
This is a product thesis, not a prediction that progress must follow an exponential curve or that one provider will win every task. Model improvements create opportunities. A platform still has to prove that those improvements help a real customer.
Native speech-to-speech is the architectural choice
OpenAI's Realtime guide describes a speech-to-speech agent that works directly with audio, maintains conversation state, and calls tools. The browser connects through WebRTC, with application code still responsible for the agent's business logic. Official OpenAI Realtime documentation.
A traditional voice pipeline turns speech into text, passes that text to a language model, then synthesizes its reply into speech. Native Realtime voice handles the spoken exchange directly, without those separate recognition and synthesis steps in the model's conversation path. Official OpenAI speech-to-speech documentation.
Native speech-to-speech modelSpoken response
For the website conversations we want to serve, we see native speech-to-speech as the foundation to build on. Our bet is that better audio understanding, spoken interaction, and tool use can expand what a useful business agent does. It is a focused product direction, not a claim that every speech pipeline is obsolete or that a model guarantees a better result for every task.
The model still needs your opening hours, services, submission rules, and a clear way to hand a request to your team. Those details belong to the business and the product around it.
A note on the current model options
Audentic currently supports Realtime model options and a configurable GPT-Live option. GPT-Live is a distinct architecture: it handles speech while delegating task work to a separate backend. Our native Realtime focus should not be read as a claim that every current configuration is the same end-to-end speech model. Official OpenAI GPT-Live documentation.
The work around the model is the product
We use “harness” to mean the application that supplies business context, connects the conversation to approved capabilities, and gives the operator visibility and control. For a website agent, the important questions are practical:
- What information can it use, and who reviews that information?
- When can a visitor start a conversation, and on which website?
- What can the assistant show or prepare during the conversation?
- What must the visitor confirm before a request is saved?
- What can the business review afterward, and what is still unknown?
That is where Audentic is concentrating its build. Centering the platform on OpenAI Realtime lets us spend more attention on making that native voice architecture work well for a business operator. It also means dependence on OpenAI's availability and model behavior, a tradeoff we accept and have to manage.
Give the operator a complete daily workflow
The current platform connects five pieces of that workflow:
- Build from the business. Describe the agent's purpose, type or talk with the builder, and review the saved instructions before publishing. Create an agent.
- Keep knowledge reviewable. Add business text, a public webpage, uploaded reference materials, or notes. Inspect the extracted text and stay within the documented context limits. Knowledge.
- Publish deliberately. Test a conversation, approve exact website origins, and copy the generated embed. Visitors choose when to start and grant microphone access. Website widget.
- Make the next step visible. Optional on-screen choices and forms can change with the conversation. A visitor must send a form or confirm a follow-up request before it reaches the owner's Inbox. Conversations and requests.
- Review and improve. Read saved transcripts, handle visitor submissions, and inspect activity and cost coverage. Connected time and historical estimates are different from verified costs. Usage and plans.
This is our foundation for a business platform. The human still owns the business facts, decides what to publish, and acts on the requests customers submit.
Focus has to produce a visible difference
Other platforms already have useful capabilities. Vapi documents OpenAI Realtime and a GPT-Live private beta alongside multiple providers. Retell has website voice and callback widgets. ElevenLabs Agents has provider choice, workflows, and web deployment. Model access and an embed button are not enough to distinguish Audentic. Vapi Realtime, Vapi GPT-Live, Retell widget, ElevenAgents overview.
Our intended difference begins with dedicated focus on native OpenAI Realtime speech-to-speech, and continues through the path from “here is my business” to a useful conversation and an understandable next action. We want a nontechnical operator to be able to correct the agent, see what happened, and decide what to do next. That is an outcome to keep testing, not a claim that competitors cannot use Realtime or serve business users.
A stronger model still needs a real-world test
Start with a narrow job: answer questions about a service and let a visitor send a request when the answer needs a human. Try the same scenarios before and after changing a model, instruction, or feature. A good update should preserve the facts and submission rules as well as improve the conversation.
Test on the host website itself. Site-builder restrictions, browser permissions, iframes, and security policies can affect the embed. A successful studio preview does not prove a working customer journey on another site. The website guide covers those constraints.
Where we are taking the platform
We want the business workflow to become easier to evaluate and operate as OpenAI voice models advance. That means better evidence about conversation outcomes, clearer handoffs, and deeper operator tools. Planned work belongs on the roadmap until it is available and tested.
Today, Audentic's documented scope is website voice agents, business knowledge, widget interactions, conversation review, and visitor requests. External bookings, outbound campaigns, and a complete CRM are beyond that scope. Paid subscriptions are coming soon; current plan details are in usage and billing.
If that starting point fits your business, follow the quickstart. If your needs are different, read the comparisons with Vapi, Retell AI, and ElevenLabs Agents before deciding.