Get in Touch

Course Outline

Introduction to Biren GPU Architecture

  • Overview of Biren technology and its primary use cases.
  • Hardware configuration: cores, memory, and compute clusters.
  • Comparative analysis with NVIDIA and AMD GPUs.

Configuring the Biren Programming Environment

  • Installation of the Biren SDK and runtime components.
  • Insights into the toolchain and compiler model.
  • Foundational project structure and build workflows.

GPU Programming Using the Biren Stack

  • Thread and block organizational models.
  • Memory management strategies and data transfer mechanisms.
  • Kernel development practices and launch patterns.

Migrating from CUDA to Biren

  • Techniques for translating CUDA codebases.
  • Mapping common APIs and necessary adaptations.
  • Laboratory sessions focused on code conversion and practice.

Debugging and Profiling Methods

  • Utilization of Biren’s native debugger and profiler tools.
  • Strategies for identifying performance bottlenecks.
  • Analysis of memory access patterns and subsequent optimization.

Advanced Optimization Techniques

  • Thread scheduling mechanisms and instruction pipelining.
  • Application of loop unrolling and shared memory utilization.
  • Sophisticated kernel tuning to maximize throughput.

Case Studies and Practical Applications

  • Model training procedures using Biren accelerators.
  • Porting and profiling exercises for vision or NLP models.
  • Performance benchmarking against CUDA/NVIDIA standards.

Summary and Future Directions

Requirements

  • A solid grasp of GPU architecture and parallel processing concepts.
  • Hands-on experience with CUDA, OpenCL, or comparable GPU programming frameworks.
  • Proficiency with deep learning frameworks such as PyTorch or TensorFlow.

Target Audience

  • HPC developers.
  • AI infrastructure engineers.
  • Specialists in performance optimization.
 21 Hours

Number of participants


Price per participant

Testimonials (2)

Upcoming Courses

Related Categories