GPT-4o (OpenAI)
OpenAI unveils GPT-4o, an "omni" model natively processing text, audio and image in real time, faster and cheaper than GPT-4.
Key points
- OpenAI reveals GPT-4o on May 13, 2024, an omni model natively handling text, audio and image.
- Real-time processing, reduced latency, lower cost than GPT-4.
- Fluid voice and visual interactions, close to a natural exchange.
- It accelerates the adoption of AI assistants in everyday and mobile uses.
Analysis
GPT-4o lowers cost and latency while unifying the modalities, which makes AI assistants more natural to use day to day, notably by voice. This fluidity widens the moments and contexts in which people query an AI rather than a classic engine.
For visibility, the rise of voice and multimodal interactions reinforces the importance of a single, well-formulated answer: by voice, there is often only one answer delivered, without a list of links. Being the chosen source becomes even more critical than in text search.
What to do
- Optimize for direct and concise answers to natural questions, phrased the way people ask them out loud.
- Structure factual data (hours, prices, contact details, definitions) so it can be easily delivered in a voice or synthetic context.
- Track your brand's presence in multimodal and voice answers, not only in written results.
Smooth voice and visual interactions. Accelerated everyday adoption of AI assistants.