Principal Networking AI Systems Architect
פורסם היום · 0 מועמדים
- Ph.D in electrical engineering, machine-learning, computer-science or another relevant field
- 10+ years of deep technical experience in high-performance network architecture, data center networking, or distributed systems design
- Mastery of high-speed interconnect protocols including InfiniBand and/or advanced Ethernet architectures
- Thorough experience driving high-impact projects centered on modern AI/ML such as LLMs/agents, deep-learning, black-box optimization or another relevant field
- Deep knowledge of AI/ML and networking-hardware/system-architecture
- Demonstrated track record of architecting and deploying multi-thousand-node GPU clusters for hyperscale cloud environments
- Deep knowledge of NVIDIA networking technologies, including BlueField DPUs, Quantum InfiniBand switches, and Spectrum Ethernet platforms
- Expertise in in-network computing, telemetry, adaptive routing, and telemetry-driven network optimization
- High energy and a positive, proactive and curious approach
חולץ מתיאור המשרה · מתעדכן אוטומטית
תיאור המשרה המלא
המשרה המקורית · נשמר לעיוןWe are seeking a highly skilled Principal Networking AI System Architect to join as a key contributor to the Applied Networking AI group. In this role you will scope and lead AI based solutions for networking technologies and drive their integration across teams.You’ll lead a portfolio and roadmap of projects that encompass agentic-AI for data-center management, predictive-resiliency, optimization and more. By collaborating closely with subject-matter-experts (SMEs), applied-researchers, product-managers, architects, data-engineers and other stakeholders you will push the envelope forward in using cutting-edge technologies and data-driven insights to improve NVIDIA's products.
What You'll Be Doing
• Build a shared roadmap and vision for AI based data-center management solutions spanning LLM intelligence for troubleshooting, predictive-resiliency and AIOPS, black-box optimization and performance tuning.
• Work closely with engineering and reliability teams to scope and define workflows utilizing and benefitting from AI/ML.
• Drive the integration of AI capabilities into system architecture and engineering workflows.
• Identify system-level opportunities for failure management, automated troubleshooting, performance improvement, and resource optimization.
• Translate system behavior, dependencies, data, and operational constraints into formulated research problems.
What We Need To See
• Ph.D in electrical engineering, machine-learning, computer-science or another relevant field.
• 10+ years of deep technical experience in high-performance network architecture, data center networking, or distributed systems design.
• Mastery of high-speed interconnect protocols including InfiniBand and/or advanced Ethernet architectures.
• Thorough experience driving high-impact projects centered on modern AI/ML such as LLMs/agents, deep-learning, black-box optimization or another relevant field.
• Deep knowledge of AI/ML and networking-hardware/system-architecture.
• Excellent ability to convey and communicate data-based insights to stakeholders and management.
• Experience demonstrating an excellent track of collaboration with hands-on teams.
Ways To Stand Out From The Crowd
• Demonstrated track record of architecting and deploying multi-thousand-node GPU clusters for hyperscale cloud environments.
• Deep knowledge of NVIDIA networking technologies, including BlueField DPUs, Quantum InfiniBand switches, and Spectrum Ethernet platforms.
• Expertise in in-network computing, telemetry, adaptive routing, and telemetry-driven network optimization.
• High energy and a positive, proactive and curious approach.
, , JR2024255
שאלות על המשרה
- המשרה לא ציינה שכר. אנחנו מציגים שכר רק כשהמעסיק מפרסם אותו.
- Ph.D in electrical engineering, machine-learning, computer-science or another relevant field, 10+ years of deep technical experience in high-performance network architecture, data center networking, or distributed systems design, Mastery of high-speed interconnect protocols including InfiniBand and/or advanced Ethernet architectures, Thorough experience driving high-impact projects centered on modern AI/ML such as LLMs/agents, deep-learning, black-box optimization or another relevant field, Deep knowledge of AI/ML and networking-hardware/system-architecture