HPC Support Team Lead
פורסם לפני 20 ימים · 0 מועמדים
התפקיד במילים פשוטות
התפקיד כולל ניהול של צוות תמיכה Tier 2.5, ניטור מדדי שירות ותפעול ותמיכה בתשתית ה-HPC והבינה המלאכותית המרכזית של המכון. המוביל יספק פתרונות מתקדמים לתקלות מערכתיות מורכבות ויסייע לחוקרים ולצוותים פנימיים בהרצה ואופטימיזציה של תוכנות מדעיות וספריות פייתון מתקדמות.
- Significant experience managing various Linux environments, over 10 years
- Team management experience in the HPC/AI field, 5 years
- Experience with DevOps tools such as Jenkins and an IaC approach, at least 5 years
- Experience working with compilers, MPI libraries, and scientific computing environments
- High-level proficiency in Hebrew and English, written and spoken
חולץ מתיאור המשרה · מתעדכן אוטומטית
למי זה מתאים
התפקיד מתאים לבעלי ניסיון ניהולי של לפחות 5 שנים בתחום ה-HPC/AI ומעל 10 שנות ניסיון בניהול סביבות לינוקס שונות. הוא פחות יתאים למי שאין לו ניסיון מעשי בכלי DevOps, עבודה עם תזמוני משימות כמו Slurm ופתרון בעיות ביצועי CPU/GPU.
תיאור המשרה המלא
המשרה המקורית · נשמר לעיוןYour role will include:
Responsible for managing the Tier 2.5 team, monitoring service metrics, and operating and supporting the institute's core HPC infrastructure, providing advanced technological support to institute researchers and internal teams, including among others:
Supporting High-Performance Computing (HPC) and Artificial Intelligence (AI) infrastructure.
Handling complex service tickets in the institute's core HPC/AI environments and providing advanced solutions for systemic issues.
Providing professional support to internal teams and institute researchers on all matters related to compiling, optimizing, and running complex scientific software in HPC/AI environments, including tools such as WRF, SAM, and Flash.
Working with scientific development environments and advanced Python libraries, including TensorFlow, PyTorch, Dask, mpi4py, NumPy, and SciPy.
Operating and optimizing job scheduling engines such as Slurm, LSF, and PBS.
Skills and abilities:
Significant experience managing various Linux environments, over 10 years -required.
Team management experience in the HPC/AI field, 5 years -required.
Experience working with High-Performance Computing (HPC) and/or AI environments -required.
Experience with DevOps tools such as Jenkins and an IaC approach, at least 5 years -required.
Experience working with compilers, MPI libraries, and scientific computing environments -required.
Experience troubleshooting complex issues in HPC/AI environments, including CPU/GPU performance testing.
Experience debugging scientific applications deployed across multiple servers and GPU accelerators.
Experience working with job scheduling engines such as Slurm, LSF, and PBS.
High-level proficiency in Hebrew and English, written and spoken.
Strong ability to analyze system failures and solve complex problems in production environments.
Ability to work independently, in a team, and in a dynamic, multitasking environment.
Responsibility, initiative, and strong self-learning ability.
Systems thinking and the ability to lead complex technological processes.
Analytical thinking and the ability to quickly learn new technologies.
Strong service orientation and interpersonal communication skills.
שאלות על המשרה
- המשרה לא ציינה שכר. אנחנו מציגים שכר רק כשהמעסיק מפרסם אותו.
- Significant experience managing various Linux environments, over 10 years, Team management experience in the HPC/AI field, 5 years, Experience with DevOps tools such as Jenkins and an IaC approach, at least 5 years, Experience working with compilers, MPI libraries, and scientific computing environments, High-level proficiency in Hebrew and English, written and spoken