The Rise of Multimodal Artificial Intelligence
Time:
2022-12-07
"Modality" is a biological concept proposed by the German physicist Helmholtz, that is, the channels through which organisms receive information through sensory organs and experiences, such as humans have sight, hearing, touch, taste and smell modes. Multimodal refers to the integration of multiple senses, while multimodal interaction refers to people communicating with computers through multiple channels such as voice, body language, information carriers (text, pictures, audio, video), and environment, fully simulating the interaction between people.
Traditional deep learning algorithms focus on training their models from a single data source. For example, a computer vision model is trained on a set of images, an NLP model is trained on text content, and speech processing involves the creation of acoustic models, wake word detection, and noise cancellation. This type of machine learning is related to single-modal AI, where the results are all mapped to a single source of data types. Multimodal AI is the ultimate fusion of computer vision and interactive AI intelligence models, providing calculators with scenarios closer to human perception.
The latest example of multimodal AI is OpenAI's DALL-E, which is named after the homophony of artists Salvador-Dalí and Pixar's WALL-E. It can generate corresponding images from text descriptions. For example, when the text description "a donut-shaped clock" is sent to the model, it can generate the following image.

Relevant information
2022-12-07
2022-12-07
GRIT (Shenzhen) Technology Co., Ltd.
© COPYRIGHT 2023 GRIT (Shenzhen) Technology Co., Ltd. ALL RIGHTS RESERVED 粤ICP备11072828号 Powered by www.300.cn SEO
