Research Team Lead - Distributed AI Systems & Large-Scale Infrastructure
פורסם אתמול · 0 מועמדים
תיאור המשרה המלא
המשרה המקורית · נשמר לעיוןA global technology R&D center is developing advanced research and high-level design solutions for information and communications technology infrastructure, smart devices, AI systems, and next-generation connectivity.
The company’s teams work on cutting-edge technologies across systems software, AI infrastructure, distributed computing, hardware-software co-design, communications, and large-scale performance optimization.
This is an opportunity to lead a research and engineering team at the frontier of distributed AI systems and large-scale infrastructure, working on software infrastructure for distributed training and inference on custom accelerator hardware.
Responsibilities
• Lead and grow a team of researchers and engineers working on distributed AI infrastructure and systems software.
• Architect and own software infrastructure for distributed training and inference on custom accelerator hardware.
• Drive research and development of communication libraries, runtime systems, memory management, graph execution, and synchronization at scale.
• Optimize end-to-end performance across large-scale clusters, including scale-up multi-device and scale-out multi-node environments.
• Design and implement high-performance communication backends and collective operations such as AllReduce, AllGather, and broadcast.
• Work on distributed AI workloads, including performance optimization for large-scale training and inference systems.
• Collaborate with hardware architects, compiler teams, framework engineers, and global research teams on hardware-software co-design.
• Mentor researchers and engineers, conduct performance and growth reviews, and shape the team’s technical direction and culture.
• Partner with academic institutions and open-source communities to advance research in distributed AI systems.
• Publish and present research findings at leading international conferences and technical forums.
Required Qualifications
• B.Sc. or higher in Computer Science, Computer Engineering, Electrical Engineering, or a related field
• At least 8 years of experience in systems software development, distributed systems, or AI infrastructure
• At least 3 years of experience in a managerial or team leadership role
• Deep expertise in large-scale communication systems, including collective communication, RDMA, topology-aware routing, and bandwidth optimization
• Hands-on experience developing software infrastructure for distributed training on dedicated accelerators or heterogeneous hardware such as GPU, NPU, or TPU
• Deep knowledge of runtime systems, including scheduling, execution graphs, kernel dispatch, synchronization mechanisms, and pipeline management
• Experience with large-scale memory management, including activation checkpointing, tensor offloading, rematerialization, and KV cache management
• Strong hands-on experience with C/C++ and Python
• Experience developing high-performance, production-ready code in Linux environments
• Excellent English, including the ability to present to international audiences, write technical documents, and lead cross-team technical processes
JOB NO. 500422
שאלות על המשרה
- המשרה לא ציינה שכר. אנחנו מציגים שכר רק כשהמעסיק מפרסם אותו.