About The Opportunity
On behalf of our clients, we are hiring a
GPU Engineer Team Leader to support their high-performance AI framework engineering team. Our client is an innovative AI infrastructure pioneer specializing in GPU/NPU optimizations and advanced LLM inference frameworks at the intersection of HPC and AI systems.
Joining the team as a GPU Engineer Team Leader, you will guide a dedicated team of 4-5 engineers while remaining hands-on with high-performance GPU software development. You will play a pivotal role in driving technical delivery, conducting code reviews, and optimizing single- and multi-GPU architectures for state-of-the-art AI workloads.
Key Responsibilities
- Team Leadership & Delivery: Lead, mentor, and develop a small team of GPU/HPC engineers. Plan sprints, manage workload distribution, and ensure the timely delivery of high-quality software components.
- Mentorship & Culture: Conduct regular 1:1s, support engineers career development, and foster a culture of continuous improvement and technical excellence.
- Technical Strategy: Translate high-level technical goals from senior management into actionable tasks, surface blockers early, and maintain clear communication with stakeholders.
- Hands-on Development: Write production-ready, low-level GPU kernel code using CUDA, HIP, or OpenCL for AI training and inference workloads.
- Quality Assurance: Lead code reviews, enforce coding standards, and perform deep performance profiling and memory hierarchy optimizations to solve critical architectural challenges.
Requirements
Technical skills:
- Bachelor's degree in Computer Science, Computer Engineering, or a related technical field.
- 2+ years of professional experience writing system software for GPUs.
- Strong programming proficiency in C++ and Python.
- Direct experience writing and optimizing GPU software using CUDA, HIP, or OpenCL.
- Deep knowledge of GPU memory hierarchies, including shared memory utilization, registers, coalescing, and occupancy optimization.
- Familiarity with deep learning frameworks (such as PyTorch or TensorFlow) and how they interact with underlying GPU hardware.
Nice-to-have (Preferred Qualifications)
- Experience with distributed GPU computing, multi-GPU coordination, or parallel runtime systems.
- Strong understanding of AI model architectures (e.g., attention mechanisms, matrix operations) and their impact on GPU workload design.
- Hands-on experience with performance profiling tools such as Nsight Compute, Nsight Systems, or AMD ROCm profiler.
- Active contributions to open-source GPU/HPC projects or publications at top-tier relevant conferences (PPoPP, HPDC, SC, MICRO, etc.).
Benefits
- Competitive salary package with performance bonuses.
- Premium healthcare insurance coverage.
- Opportunity to work with cutting-edge HPC, multi-GPU systems, and generative AI infrastructure.
- Clear career growth paths and continuous professional development support.
- Dynamic, open, and technical engineering work environment.