Handle and integrate multiple input/output types such as text, images, and audio simultaneously.
3 models
by DeepMind
A few-shot visual-language model that processes interleaved sequences of images and text.
Not yet ratedby OpenAI
Aligns images and text with contrastive learning, enabling powerful zero-shot vision-language tasks.
Not yet ratedby OpenAI
A multimodal model that understands and generates text, audio, and images together.
Not yet rated