HPC Support Team Lead - 3926
פורסם אתמול · 0 מועמדים
- 10+ years of significant experience managing various Linux environments
- 5+ years of team management experience in HPC/AI environments
- At least 5 years of experience with DevOps tools such as Jenkins and Infrastructure as Code (IaC) methodologies
- Experience working with compilers, MPI libraries, and scientific computing environments
- Experience troubleshooting complex issues in HPC/AI environments, including CPU/GPU performance benchmarking and analysis
חולץ מתיאור המשרה · מתעדכן אוטומטית
תיאור המשרה המלא
המשרה המקורית · נשמר לעיוןHPC Support Team Lead
A large scientific organization in Israel’s Shephelah region is seeking an HPC Support Team Lead.
Responsibilities
• Lead and manage a Tier 2.5 support team, monitor service KPIs, and oversee the operation and support of the institute’s core HPC infrastructure.
• Provide advanced technical support to the institute’s researchers and internal teams, including:
• Support for High-Performance Computing (HPC) and Artificial Intelligence (AI) infrastructure.
• Handle complex support tickets in the institute’s core HPC/AI environments and provide advanced solutions to systemic issues.
• Provide professional support to internal teams and researchers in the compilation, optimization, and execution of complex scientific applications in HPC/AI environments, including tools such as WRF, SAM, and Flash.
• Work with scientific development environments and advanced Python libraries, including TensorFlow, PyTorch, Dask, mpi4py, NumPy, and SciPy.
• Operate and optimize job scheduling systems such as Slurm, LSF, and PBS.
Qualifications & Skills
• 10+ years of significant experience managing various Linux environments – mandatory.
• 5+ years of team management experience in HPC/AI environments – mandatory.
• Proven experience working with High-Performance Computing (HPC) and/or AI environments – mandatory.
• At least 5 years of experience with DevOps tools such as Jenkins and Infrastructure as Code (IaC) methodologies – mandatory.
• Experience working with compilers, MPI libraries, and scientific computing environments – mandatory.
• Experience troubleshooting complex issues in HPC/AI environments, including CPU/GPU performance benchmarking and analysis.
• Experience debugging scientific applications deployed across large numbers of servers and GPU accelerators.
• Experience working with job scheduling systems such as Slurm, LSF, and PBS.
• High-level proficiency in Hebrew and English, both written and spoken.
• Strong analytical and troubleshooting skills, with the ability to diagnose systemic issues and solve complex problems in production environments.
• Ability to work independently and as part of a team in a dynamic, multi-task environment.
• High level of responsibility, initiative, and self-learning ability.
• Strong systems-level perspective and the ability to lead complex technological processes.
• Analytical thinking and the ability to quickly learn and adopt new technologies.
• Strong service orientation and excellent interpersonal and communication skills.
שאלות על המשרה
- המשרה לא ציינה שכר. אנחנו מציגים שכר רק כשהמעסיק מפרסם אותו.
- 10+ years of significant experience managing various Linux environments, 5+ years of team management experience in HPC/AI environments, At least 5 years of experience with DevOps tools such as Jenkins and Infrastructure as Code (IaC) methodologies, Experience working with compilers, MPI libraries, and scientific computing environments, Experience troubleshooting complex issues in HPC/AI environments, including CPU/GPU performance benchmarking and analysis