GPU Senior Software Engineer
פורסם לפני 4 ימים · 0 מועמדים
התפקיד במילים פשוטות
תפקיד זה כולל פיתוח, אופטימיזציה ותחזוקה של רכיבי תוכנה עבור מעבדי NPU/GPU בסביבות AI ו-ML מבוזרות. העבודה היום-יומית כוללת תכנון מערכות ניהול זיכרון, שיפור ביצועים, ועבודה בצמוד עם צוותי חומרה, רשת ו-AI. בנוסף, התפקיד כולל כתיבת קוד ב-C++ ו-Python וביצוע סקירות קוד.
- 5+ years of experience in distributed system programming
- 3+ years of experience with NPU programming (Triton, CUDA, HIP, OpenCL)
- Expert-level C/C++ programming with focus on performance optimization
- Expert-level Python programming with focus on DL/ML frameworks (PyTorch/JAX/etc)
- Deep understanding of NPU architecture, memory tiering, and programming models
- Experience with NPU software stack development
- Experience with large-scale NPU systems (100+ NPUs)
- Experience with DL/ML workloads (oriented AI) and distributed training / inferencing
- Familiarity with containerization and orchestration
חולץ מתיאור המשרה · מתעדכן אוטומטית
למי זה מתאים
התפקיד מתאים למהנדסי תוכנה מנוסים עם לפחות 5 שנות ניסיון במערכות מבוזרות ו-3 שנות ניסיון בתכנות NPU/CUDA, בעלי שליטה מעולה ב-C++ ו-Python. הוא פחות מתאים לבעלי ניסיון מועט בתכנות Low-Level או ללא רקע במערכות AI/ML מבוזרות.
תיאור המשרה המלא
המשרה המקורית · נשמר לעיוןLocation: Tel Aviv #Hybrid DriveNets is a leader in high-scale disaggregated networking solutions. Founded in 2015, DriveNets modernizes the way service providers, cloud providers and hyperscalers build networks. Supporting the largest network in the world, more than half of AT&T’s backbone traffic is running on DriveNets’ Network Cloud open disaggregated architecture. Raising $587 million in three funding rounds, DriveNets is disrupting the networking market from high-scale architecture to AI platforms, and is bringing onboard the most talented people. We are seeking people that want to make an impact on the world’s leading communication networks and are experienced in networking architecture or AI infrastructure solutions. Job Summary We are seeking a skilled software engineer to join our NPU software stack development team. This role involves developing high-performance GPU programming frameworks, runtime systems, and libraries for AI/ML workloads. You will be responsible for implementing, optimizing, and maintaining GPU software stack components to support distributed AI training and inference. Key Responsibilities • Identify bottlenecks, analysis and optimize in distributed NPU eco-system • Design and develop NPU memory management system • Design and develop optimized NPU development framework, execution path and debugging • Develop compatibility with AI frameworks (Triton, PyTorch, JAX) • Write high-quality, well-tested code with comprehensive documentation • Collaborate with other teams (Hardware, Network, QA, AI Framework Integration) • Participate in code reviews and technical design discussions
Requirements: Required Qualifications • 5+ years of experience in distributed system programming • 3+ years of experience with NPU programming (Triton, CUDA, HIP, OpenCL) • Expert-level C/C++ programming with focus on performance optimization • Expert-level Python programming with focus on DL/ML frameworks (PyTorch/JAX/etc) • Deep understanding of NPU architecture, memory tiering, and programming models • Knowledge of NPU runtime systems • Experience with performance profiling and optimization tools • Strong problem-solving and debugging skills • Experience with version control systems, Ticking system and collaborative development • Team player with excellent communication skills • Fast learner, highly organized, detail-oriented with high motivation Preferred Qualifications • Experience with NPU software stack development • Experience with large-scale NPU systems (100+ NPUs) • Experience with DL/ML workloads (oriented AI) and distributed training / inferencing • Familiarity with containerization and orchestration
שאלות על המשרה
- המשרה לא ציינה שכר. אנחנו מציגים שכר רק כשהמעסיק מפרסם אותו.
- היברידי
- 5+ years of experience in distributed system programming, 3+ years of experience with NPU programming (Triton, CUDA, HIP, OpenCL), Expert-level C/C++ programming with focus on performance optimization, Expert-level Python programming with focus on DL/ML frameworks (PyTorch/JAX/etc), Deep understanding of NPU architecture, memory tiering, and programming models