Research Team Lead – Distributed AI Systems & Large-Scale Infrastructure
פורסם לפני 30 ימים · 0 מועמדים
התפקיד במילים פשוטות
תפקיד מוביל צוות מחקר המתמקד בתשתיות בינה מלאכותית בקנה מידה גדול ומערכות מבוזרות. התפקיד כולל הגדרת חזון טכנולוגי, ניהול פרויקטים מקצה לקצה ופיתוח תוכנה בתשתיות אימון מבוזר בחומרה הטארוגנית עם דגש על C/C++ ו-Python בסביבת Linux.
- B.Sc. or higher in Computer Science, Computer Engineering, Electrical Engineering, or a closely related field
- 8+ years of experience in systems software, distributed computing, or AI infrastructure
- 3+ years in a leadership or team lead role
- Deep expertise in large-scale communication systems: collective communication, RDMA, network topology-aware routing, and bandwidth optimization
- Hands-on experience building software infrastructure for distributed training on custom accelerators or heterogeneous hardware (GPU, NPU, TPU)
- M.Sc. or Ph.D. in a relevant field, with a strong publication record at systems or ML venues (EuroSys, OSDI, SC, NeurIPS, MLSys, ISCA)
- Hands-on experience with communication frameworks such as NCCL, MPI, HCCL, or UCX
- Experience with compiler and graph optimization for AI workloads (XLA, TVM, Triton, or custom operator fusion)
- Background in mixed-precision training, model parallelism (Tensor Parallelism, Pipeline Parallelism, Expert Parallelism), and large model co-design
- Experience profiling and debugging performance bottlenecks on heterogeneous clusters using tools like Chrome tracing, nsight, or custom profilers
חולץ מתיאור המשרה · מתעדכן אוטומטית
למי זה מתאים
התפקיד מתאים למועמדים בעלי תואר ראשון ומעלה במדעי המחשב או הנדסה, עם מעל 8 שנות ניסיון במערכות תוכנה וחישוב מבוזר ו-3 שנות ניסיון לפחות בהובלת צוות. הוא פחות מתאים למי שחסר ניסיון הובלתי או רקע מעמיק בתקשורת בקנה מידה גדול ובניהול זיכרון מבוזר.
תיאור המשרה המלא
המשרה המקורית · נשמר לעיוןRequirements
• B.Sc. or higher in Computer Science, Computer Engineering, Electrical Engineering, or a closely related field
• 8+ years of experience in systems software, distributed computing, or AI infrastructure, with 3+ years in a leadership or team lead role
• Deep expertise in large-scale communication systems: collective communication, RDMA, network topology-aware routing, and bandwidth optimization
• Hands-on experience building software infrastructure for distributed training on custom accelerators or heterogeneous hardware (GPU, NPU, TPU)
• Strong knowledge of runtime systems: scheduling, execution graphs, kernel dispatch, synchronization primitives, and pipeline management
• Experience with memory management at scale: activation checkpointing, tensor offloading, rematerialization, KV cache management
• Proficiency in C/C++ and Python, with a focus on high-performance, production-quality code in Linux environments
• Proven ability to define technical vision, lead multi-person projects end-to-end, and deliver results under research and engineering timelines
• Excellent communication skills in English — confident presenting to international audiences, writing technical reports, and driving cross-team alignment
• Strong collaborative mindset and experience working in globally distributed, multicultural teams
Ways to Stand Out From the Crowd
M.Sc. or Ph.D. in a relevant field, with a strong publication record at systems or ML venues (EuroSys, OSDI, SC, NeurIPS, MLSys, ISCA)
• Hands-on experience with communication frameworks such as NCCL, MPI, HCCL, or UCX
• Experience with compiler and graph optimization for AI workloads (XLA, TVM, Triton, or custom operator fusion)
• Background in mixed-precision training, model parallelism (Tensor Parallelism, Pipeline Parallelism, Expert Parallelism), and large model co-design
• Experience profiling and debugging performance bottlenecks on heterogeneous clusters using tools like Chrome tracing, nsight, or custom profilers
שאלות על המשרה
- המשרה לא ציינה שכר. אנחנו מציגים שכר רק כשהמעסיק מפרסם אותו.
- B.Sc. or higher in Computer Science, Computer Engineering, Electrical Engineering, or a closely related field, 8+ years of experience in systems software, distributed computing, or AI infrastructure, 3+ years in a leadership or team lead role, Deep expertise in large-scale communication systems: collective communication, RDMA, network topology-aware routing, and bandwidth optimization, Hands-on experience building software infrastructure for distributed training on custom accelerators or heterogeneous hardware (GPU, NPU, TPU)