GPU AI Performance Engineer
Qualcomm
We are looking for a GPU AI Performance Engineer to help design, optimize, and accelerate AI workloads on GPU-based platforms. This role focuses on improving the performance, memory efficiency, and scalability of modern AI models — especially LLMs, transformer workloads, MoE models, and computer vision pipelines — by working across the stack from model architecture to compiler/runtime integration and low-level kernel optimization.
Responsibilities
- Analyze and evaluate GPU architecture/microarchitecture for performance optimizations of AI workloads
- Work across the AI stack, from model graphs and inference runtimes down to GPU kernels and compiler IR
- Collaborate with hardware, software, and ML teams to identify performance bottlenecks and propose architectural or algorithmic optimizations
- Analyze AI workload characteristics and correlate their behavior across different GPU generations
- Contribute to architectural trade-off studies and influence GPU roadmap decisions with data-driven insights
Preferred Skills:
- Strong understanding of CPU/GPU architecture
- Experience in Python, C++, and ML frameworks (e.g., TensorFlow, PyTorch)
- Skills: C/C++ Programming Language, Scripting (Python/Perl), Assembly, Verilog/SystemVerilog
- Familiarity with AI inference runtimes and deployment stacks for model compilation, optimization, and execution on GPU/accelerator platforms
- Good understanding of common neural network layers and operations, including what they do and how they affect model behavior and performance
Nice to have:
- Experience with LLM inference engines such as llama.cpp
- Exposure to Triton, TTIR, TTGIR, or GPU compiler pipelines
- Knowledge of quantization formats and tradeoffs
- Understanding of MoE architectures
- Experience with GPU driver and compiler development
- Experience with OpenCL or Cuda development
Don't want to miss the next one?
Subscribe to daily email alerts for roles matching your interests.