HPC Support Team Lead
Posted 20 days ago · 0 applicants
Saving, applying or scoring takes a few seconds to set up your free account.
The role in plain words
- Significant experience managing various Linux environments, over 10 years
- Team management experience in the HPC/AI field, 5 years
- Experience with DevOps tools such as Jenkins and an IaC approach, at least 5 years
- Experience working with compilers, MPI libraries, and scientific computing environments
- High-level proficiency in Hebrew and English, written and spoken
Extracted from the job description · kept up to date automatically
Who this suits
Full job description
Original listing · kept for referenceYour role will include:
Responsible for managing the Tier 2.5 team, monitoring service metrics, and operating and supporting the institute's core HPC infrastructure, providing advanced technological support to institute researchers and internal teams, including among others:
Supporting High-Performance Computing (HPC) and Artificial Intelligence (AI) infrastructure.
Handling complex service tickets in the institute's core HPC/AI environments and providing advanced solutions for systemic issues.
Providing professional support to internal teams and institute researchers on all matters related to compiling, optimizing, and running complex scientific software in HPC/AI environments, including tools such as WRF, SAM, and Flash.
Working with scientific development environments and advanced Python libraries, including TensorFlow, PyTorch, Dask, mpi4py, NumPy, and SciPy.
Operating and optimizing job scheduling engines such as Slurm, LSF, and PBS.
Skills and abilities:
Significant experience managing various Linux environments, over 10 years -required.
Team management experience in the HPC/AI field, 5 years -required.
Experience working with High-Performance Computing (HPC) and/or AI environments -required.
Experience with DevOps tools such as Jenkins and an IaC approach, at least 5 years -required.
Experience working with compilers, MPI libraries, and scientific computing environments -required.
Experience troubleshooting complex issues in HPC/AI environments, including CPU/GPU performance testing.
Experience debugging scientific applications deployed across multiple servers and GPU accelerators.
Experience working with job scheduling engines such as Slurm, LSF, and PBS.
High-level proficiency in Hebrew and English, written and spoken.
Strong ability to analyze system failures and solve complex problems in production environments.
Ability to work independently, in a team, and in a dynamic, multitasking environment.
Responsibility, initiative, and strong self-learning ability.
Systems thinking and the ability to lead complex technological processes.
Analytical thinking and the ability to quickly learn new technologies.
Strong service orientation and interpersonal communication skills.
Questions about this role
- This listing did not state a salary. We only show pay when the employer publishes it.
- Significant experience managing various Linux environments, over 10 years, Team management experience in the HPC/AI field, 5 years, Experience with DevOps tools such as Jenkins and an IaC approach, at least 5 years, Experience working with compilers, MPI libraries, and scientific computing environments, High-level proficiency in Hebrew and English, written and spoken