News
Google releases its first native multimodal embedding model, Gemini Embedding 2.
Google has released its first native multimodal embedding model, Gemini Embedding 2, which supports mapping text, images, videos, audio, and documents to the same embedding space and can recognize semantic intent in 100 languages. The model can process up to 6 images, 120 seconds of video, 6 pages of PDF, and direct audio input in a single request, making it suitable for scenarios such as RAG, semantic search, sentiment analysis, and data clustering.