Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Core Performance Concepts and Metrics
- Analysis of latency, throughput, power consumption, and resource usage
- Distinguishing between system-level and model-level bottlenecks
- Profiling techniques for inference versus training contexts
Profiling with Huawei Ascend
- Leveraging CANN Profiler and MindInsight
- Diagnosing kernel and operator performance
- Understanding offload patterns and memory mapping
Profiling on Biren GPU
- Exploring Biren SDK performance monitoring capabilities
- Managing kernel fusion, memory alignment, and execution queues
- Conducting power and temperature-aware profiling
Profiling on Cambricon MLU
- Utilizing BANGPy and Neuware performance tools
- Gaining kernel-level visibility and interpreting logs
- Integrating the MLU profiler with deployment frameworks
Graph and Model-Level Optimization
- Strategies for graph pruning and quantization
- Techniques for operator fusion and computational graph restructuring
- Standardizing input sizes and tuning batch processing
Memory and Kernel Optimization
- Improving memory layout and data reuse
- Managing buffers efficiently across different chipsets
- Applying kernel-level tuning techniques specific to each platform
Cross-Platform Best Practices
- Ensuring performance portability through abstraction strategies
- Developing shared tuning pipelines for multi-chip environments
- Case study: optimizing an object detection model across Ascend, Biren, and MLU
Conclusion and Future Directions
Requirements
- Practical experience in AI model training or deployment pipelines
- A solid grasp of GPU/MLU compute principles and model optimization
- Fundamental knowledge of performance profiling tools and key metrics
Target Audience
- Performance engineers
- Machine learning infrastructure teams
- AI system architects
21 Hours