Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction to Multimodal LLMs in Vertex AI
- Exploring the multimodal capabilities offered by Vertex AI
- An overview of Gemini models and their supported modalities
- Examining enterprise and research use cases
Establishing the Development Environment
- Configuring Vertex AI to support multimodal workflows
- Managing datasets across different modalities
- Hands-on lab: Setting up the environment and preparing datasets
Long Context Windows and Advanced Reasoning
- Gaining insight into long-context workflow mechanics
- Applying concepts to planning and decision-making scenarios
- Hands-on lab: Executing long-context analysis
Cross-Modal Workflow Design
- Synthesizing text, audio, and image analytics
- Sequencing multimodal steps within pipeline architectures
- Hands-on lab: Architecting a multimodal pipeline
Managing Gemini API Parameters
- Setting up multimodal inputs and outputs
- Enhancing inference speed and operational efficiency
- Hands-on lab: Fine-tuning Gemini API parameters
Advanced Applications and Integrations
- Developing interactive multimodal agents and assistants
- Connecting external APIs and third-party tools
- Hands-on lab: Developing a comprehensive multimodal application
Evaluation and Iterative Improvement
- Assessing multimodal system performance
- Utilizing metrics for accuracy, alignment, and drift detection
- Hands-on lab: Evaluating the effectiveness of multimodal workflows
Conclusion and Recommended Next Steps
Requirements
- Solid proficiency in Python programming
- Practical experience in developing machine learning models
- Working knowledge of multimodal data types, including text, audio, and images
Target Audience
- AI researchers
- Senior developers
- Machine learning scientists
14 Hours