Host and Systems Performance Senior Manager
Posted 11 days ago · 0 applicants
Saving, applying or scoring takes a few seconds to set up your free account.
The role in plain words
You will lead a team of engineers focused on measuring, analyzing, and improving the performance of network interface cards, host systems, and networking stacks across NVIDIA products. Your day-to-day work involves defining performance test plans, guiding pre-silicon and post-silicon analysis, and identifying system-level bottlenecks. You will also collaborate across hardware, software, and architecture teams to resolve performance issues and align on product goals.
- B.Sc. or M.Sc. in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience
- Experience in performance analysis, systems engineering, networking, or HPC/AI infrastructure
- 10+ years of software engineering experience
- 4+ years in a leadership or management role
- Experience with high-performance networking technologies such as RDMA, RoCE, InfiniBand, Ethernet, MPI or NCCL
- C/C++ experience
- Experience with NICs or DPU architecture, networking offloads, host datapath behavior, networking services, or DPU offloads
- Experience with AI/HPC cluster performance, distributed training, inference workloads, collective communication libraries, telemetry pipelines, benchmark automation, CUDA, NCCL internals
- Linux kernel networking, DPDK, OVS, storage acceleration, or security offloads
Extracted from the job description · kept up to date automatically
Who this suits
This role suits experienced engineering leaders with over ten years of software engineering experience, including at least four years in management, and a strong background in high-performance networking technologies like RDMA, InfiniBand, or Ethernet. It is less ideal for those who lack hands-on experience with system performance profiling, benchmarking, and host architectures.
Full job description
Original listing · kept for referenceWe are looking for a Host and Systems Performance Manager to join NVIDIA’s Networking Performance team!
In this role, you will lead a team of engineers who measure, analyze, and improve NIC, host, DOCA networking stack, and system-level performance across NVIDIA networking products and platforms. We work across the full product lifecycle, from early architecture and pre-silicon modeling through post-silicon bring-up, validation, characterization, and GA readiness.
We are looking for someone who enjoys building teams, turning complex performance data into clear direction, and partnering across engineering groups to improve products that support AI, HPC, cloud, and accelerated networking workloads.
What You’ll Be Doing
• Lead, coach, and develop a team focused on NIC, host, DOCA networking stack, and system-level performance. You will define performance test plans, methodologies, metrics, and success criteria for new NVIDIA networking technologies.
• Guide pre-silicon and post-silicon performance planning, analysis, reporting, and readiness reviews. Your team will benchmark and profile workloads across RDMA, RoCE, InfiniBand, Ethernet, MPI, NCCL, DOCA, storage, security, and host networking stacks.
• Identify bottlenecks across NIC, DPU, CPU, memory, PCIe, firmware, drivers, Linux networking, DOCA, and full-system architecture. We will look to you to lead root-cause analysis, coordinate mitigation plans, and help hardware, firmware, software, architecture, validation, and product teams align on performance goals.
What We Need To See
• B.Sc. or M.Sc. in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
• Experience in performance analysis, systems engineering, networking, or HPC/AI infrastructure.
• 10+ years of software engineering experience, including 4+ years in a leadership or management role.
• Experience with high-performance networking technologies such as RDMA, RoCE, InfiniBand, Ethernet, MPI or NCCL.
• Hands-on experience with system performance analysis, benchmarking, profiling, and root-cause analysis.
• Understanding of host architecture, including CPUs, memory hierarchy, NUMA, PCIe, Linux OS, drivers, firmware, and DPU/NIC interactions.
• Experience creating performance test plans for pre-silicon or post-silicon phases.
• Programming or scripting experience with Python and Bash; C/C++ experience is helpful.
• Ability to communicate technical findings clearly and work across engineering teams.
Ways To Stand Out from the Crowd
• Experience with NICs or DPU architecture, networking offloads, host datapath behavior, networking services, or DPU offloads.
• Experience with AI/HPC cluster performance, distributed training, inference workloads, collective communication libraries, telemetry pipelines, benchmark automation, CUDA, NCCL internals.
• Linux kernel networking, DPDK, OVS, storage acceleration, or security offloads can also help you succeed in this role.
We support our people with competitive compensation, health and wellness programs, financial benefits, time away from work, family support, learning opportunities, and resources that help employees do their best work. Benefits may vary by location, role, and eligibility.
NVIDIA is an equal opportunity employer. We are committed to creating an inclusive environment for all employees and applicants. We consider qualified applicants without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, disability, age, veteran status, or other protected status.
We provide reasonable accommodations for applicants and employees with disabilities. If you need support during the application or interview process, please let our recruiting team know.
, , JR2020481
Questions about this role
- This listing did not state a salary. We only show pay when the employer publishes it.
- B.Sc. or M.Sc. in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience, Experience in performance analysis, systems engineering, networking, or HPC/AI infrastructure, 10+ years of software engineering experience, 4+ years in a leadership or management role, Experience with high-performance networking technologies such as RDMA, RoCE, InfiniBand, Ethernet, MPI or NCCL