Nvidia Corporation
Santa Clara, CA
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop silicon-measured kernel benchmarking infrastructure, model-level performance projection tooling, and agentic optimization systems that improve GPU kernels at the assembly layer. Our team works closely with compiler, kernel, hardware, and framework organizations across NVIDIA to surface bottlenecks and ship measurable gains. If driving GPU performance at the frontier of LLM inference sounds like your kind of challenge, we'd love to meet you! What you'll be doing: The role drives three interconnected systems, all aimed at accelerating NVIDIA's LLM inference stack. The first is GPU kernel microbenchmarking: measuring competing kernel implementations at real-silicon fidelity across the full configuration space that...