Learning Blog

OpenAI Announces GPT-4o


0b2b2bde6063336becccaee3572429e3f7ae666e1f06eb79aa411bae9d703c92.jpg

OpenAI has announced GPT-4o, its new flagship model that can reason across audio, vision, and text in real time.

GPT-4o (“o” for “omni”) is a step towards a more natural human-computer interaction — it accepts any combination of text, audio, image, and video as input and generates any combination of text, audio, and image outputs. It can respond to audio inputs in as little as 232 milliseconds, with an average of 320 milliseconds, which is similar to human response time(opens in a new window) in a conversation. It matches GPT-4 Turbo performance on text in English and code, with significant improvement on text in non-English languages, while also being much faster and 50% cheaper in the API. GPT-4o is especially better at vision and audio understanding compared to existing models.

How do you rate this article?

3


Learning Pages
Learning Pages

Learning Pages publishes articles about life, learning, writing and technology.


Learning Blog
Learning Blog

Posts about life, learning, writing and technology.

Publish0x Publish0x

Reward the author with $0.01 in crypto, and earn yourself as you read!

20% to author / 80% to me.
Rewards are FREE. Publish0x pays them, not you.

Page not displaying correctly?