Research Team Lead - Distributed AI Systems & Large-Scale Infrastructure
פורסם לפני 28 ימים · 0 מועמדים
התפקיד במילים פשוטות
התפקיד כולל הובלת צוות מחקר והנדסה המפתח תשתיות תוכנה לאימון והרצת מודלי AI מבוזרים על גבי חומרת מאיצים ייעודית. ביומיום התפקיד כולל תכנון ספריות תקשורת ומערכות Runtime, אופטימיזציית ביצועים בקלאסטרים גדולים (Scale-up ו-Scale-out), שיתוף פעולה עם צוותי חומרה וקומפיילרים, ופרסום ממצאי מחקר בכנסים בינלאומיים.
- B.Sc. or higher in Computer Science, Computer Engineering, Electrical Engineering, or a related field
- At least 8 years of experience in systems software development, distributed systems, or AI infrastructure
- At least 3 years of experience in a managerial or team leadership role
- Deep expertise in large-scale communication systems, including collective communication, RDMA, topology-aware routing, and bandwidth optimization
- Hands-on experience developing software infrastructure for distributed training on dedicated accelerators or heterogeneous hardware (GPU, NPU, TPU)
חולץ מתיאור המשרה · מתעדכן אוטומטית
למי זה מתאים
מתאים למובילי צוותים מנוסים עם רקע עמוק במערכות תוכנה מבוזרות, תשתיות AI, תכנות ב-C/C++ ו-Python, ומומחיות בתקשורת בקנה מידה גדול ומאיצי חומרה. פחות מתאים למי שחסר ניסיון ניהולי או ניסיון בפיתוח Low-level ותשתיות מחשוב מבוזר.
תיאור המשרה המלא
המשרה המקורית · נשמר לעיוןA global technology R&D center is developing advanced research and high-level design solutions for information and communications technology infrastructure, smart devices, AI systems, and next-generation connectivity.
The company’s teams work on cutting-edge technologies across systems software, AI infrastructure, distributed computing, hardware-software co-design, communications, and large-scale performance optimization.
This is an opportunity to lead a research and engineering team at the frontier of distributed AI systems and large-scale infrastructure, working on software infrastructure for distributed training and inference on custom accelerator hardware.
Responsibilities
• Lead and grow a team of researchers and engineers working on distributed AI infrastructure and systems software.
• Architect and own software infrastructure for distributed training and inference on custom accelerator hardware.
• Drive research and development of communication libraries, runtime systems, memory management, graph execution, and synchronization at scale.
• Optimize end-to-end performance across large-scale clusters, including scale-up multi-device and scale-out multi-node environments.
• Design and implement high-performance communication backends and collective operations such as AllReduce, AllGather, and broadcast.
• Work on distributed AI workloads, including performance optimization for large-scale training and inference systems.
• Collaborate with hardware architects, compiler teams, framework engineers, and global research teams on hardware-software co-design.
• Mentor researchers and engineers, conduct performance and growth reviews, and shape the team’s technical direction and culture.
• Partner with academic institutions and open-source communities to advance research in distributed AI systems.
• Publish and present research findings at leading international conferences and technical forums.
Required Qualifications
• B.Sc. or higher in Computer Science, Computer Engineering, Electrical Engineering, or a related field
• At least 8 years of experience in systems software development, distributed systems, or AI infrastructure
• At least 3 years of experience in a managerial or team leadership role
• Deep expertise in large-scale communication systems, including collective communication, RDMA, topology-aware routing, and bandwidth optimization
• Hands-on experience developing software infrastructure for distributed training on dedicated accelerators or heterogeneous hardware such as GPU, NPU, or TPU
• Deep knowledge of runtime systems, including scheduling, execution graphs, kernel dispatch, synchronization mechanisms, and pipeline management
• Experience with large-scale memory management, including activation checkpointing, tensor offloading, rematerialization, and KV cache management
• Strong hands-on experience with C/C++ and Python
• Experience developing high-performance, production-ready code in Linux environments
• Excellent English, including the ability to present to international audiences, write technical documents, and lead cross-team technical processes
JOB NO. 500422
שאלות על המשרה
- המשרה לא ציינה שכר. אנחנו מציגים שכר רק כשהמעסיק מפרסם אותו.
- B.Sc. or higher in Computer Science, Computer Engineering, Electrical Engineering, or a related field, At least 8 years of experience in systems software development, distributed systems, or AI infrastructure, At least 3 years of experience in a managerial or team leadership role, Deep expertise in large-scale communication systems, including collective communication, RDMA, topology-aware routing, and bandwidth optimization, Hands-on experience developing software infrastructure for distributed training on dedicated accelerators or heterogeneous hardware (GPU, NPU, TPU)