Get in Touch

Course Outline

Introduction to Multimodal LLMs in Vertex AI

  • Exploring the multimodal capabilities offered by Vertex AI
  • An overview of Gemini models and their supported modalities
  • Examining enterprise and research use cases

Establishing the Development Environment

  • Configuring Vertex AI to support multimodal workflows
  • Managing datasets across different modalities
  • Hands-on lab: Setting up the environment and preparing datasets

Long Context Windows and Advanced Reasoning

  • Gaining insight into long-context workflow mechanics
  • Applying concepts to planning and decision-making scenarios
  • Hands-on lab: Executing long-context analysis

Cross-Modal Workflow Design

  • Synthesizing text, audio, and image analytics
  • Sequencing multimodal steps within pipeline architectures
  • Hands-on lab: Architecting a multimodal pipeline

Managing Gemini API Parameters

  • Setting up multimodal inputs and outputs
  • Enhancing inference speed and operational efficiency
  • Hands-on lab: Fine-tuning Gemini API parameters

Advanced Applications and Integrations

  • Developing interactive multimodal agents and assistants
  • Connecting external APIs and third-party tools
  • Hands-on lab: Developing a comprehensive multimodal application

Evaluation and Iterative Improvement

  • Assessing multimodal system performance
  • Utilizing metrics for accuracy, alignment, and drift detection
  • Hands-on lab: Evaluating the effectiveness of multimodal workflows

Conclusion and Recommended Next Steps

Requirements

  • Solid proficiency in Python programming
  • Practical experience in developing machine learning models
  • Working knowledge of multimodal data types, including text, audio, and images

Target Audience

  • AI researchers
  • Senior developers
  • Machine learning scientists
 14 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories