דלג לתוכן הראשי

Research Team Lead – Distributed AI Systems & Large-Scale Infrastructure

Ethosiaהוד השרון, מחוז המרכז, ישראללא צויןFull-timeדרגה: לא צוין

פורסם אתמול · 0 מועמדים

שכר לא צוין במשרה זו

שמירה, הגשה או בדיקת התאמה — כמה שניות להקמת חשבון חינם.

תובנת Willbi

התפקיד במילים פשוטות

חובה
    יתרון

      חולץ מתיאור המשרה · מתעדכן אוטומטית

      למי זה מתאים

      תיאור המשרה המלא

      המשרה המקורית · נשמר לעיון

      Requirements

      • B.Sc. or higher in Computer Science, Computer Engineering, Electrical Engineering, or a closely related field

      • 8+ years of experience in systems software, distributed computing, or AI infrastructure, with 3+ years in a leadership or team lead role

      • Deep expertise in large-scale communication systems: collective communication, RDMA, network topology-aware routing, and bandwidth optimization

      • Hands-on experience building software infrastructure for distributed training on custom accelerators or heterogeneous hardware (GPU, NPU, TPU)

      • Strong knowledge of runtime systems: scheduling, execution graphs, kernel dispatch, synchronization primitives, and pipeline management

      • Experience with memory management at scale: activation checkpointing, tensor offloading, rematerialization, KV cache management

      • Proficiency in C/C++ and Python, with a focus on high-performance, production-quality code in Linux environments

      • Proven ability to define technical vision, lead multi-person projects end-to-end, and deliver results under research and engineering timelines

      • Excellent communication skills in English — confident presenting to international audiences, writing technical reports, and driving cross-team alignment

      • Strong collaborative mindset and experience working in globally distributed, multicultural teams

      Ways to Stand Out From the Crowd

      M.Sc. or Ph.D. in a relevant field, with a strong publication record at systems or ML venues (EuroSys, OSDI, SC, NeurIPS, MLSys, ISCA)

      • Hands-on experience with communication frameworks such as NCCL, MPI, HCCL, or UCX

      • Experience with compiler and graph optimization for AI workloads (XLA, TVM, Triton, or custom operator fusion)

      • Background in mixed-precision training, model parallelism (Tensor Parallelism, Pipeline Parallelism, Expert Parallelism), and large model co-design

      • Experience profiling and debugging performance bottlenecks on heterogeneous clusters using tools like Chrome tracing, nsight, or custom profilers

      אודות Ethosia
      פרופיל החברה · בקרוב

      ביקורות עובדים · בקרובעוד משרות ב-Ethosia

      שאלות על המשרה

      • המשרה לא ציינה שכר. אנחנו מציגים שכר רק כשהמעסיק מפרסם אותו.
      דומות וקשורות
      Ethosia
      פורסם אתמול · 0 מועמדים
      בדקו את ההתאמה