Head of AI Infrastructure (MLOps)
Posted 27 days ago · 0 applicants
Saving, applying or scoring takes a few seconds to set up your free account.
The role in plain words
- Proven experience as an MLOps Engineer, ML Engineer, or Data Engineer with strong ML orientation
- Strong Python skills
- hands-on experience with ML/DL frameworks (PyTorch / TensorFlow)
- Experience building model training and inference pipelines
- Experience with experiment tracking and model management tools (MLflow / Weights & Biases)
- Experience working with satellite data (EO / SAR)
- Familiarity with Ray, Airflow, or Prefect
- Experience with GPU workloads and model training optimization
- Knowledge of observability tools (Prometheus, Grafana, OpenTelemetry)
- Experience with data versioning (DVC) or feature stores
Extracted from the job description · kept up to date automatically
Who this suits
Full job description
Original listing · kept for referenceHead of MLOps & AI Infrastructure
This role bridges research and production, owning the full model lifecycle — from experimentation to operational deployment.
The position includes building scalable MLOps infrastructure, automating training and inference workflows, and continuously improving performance and reliability in distributed systems.
Key Responsibilities
• Design and lead end-to-end MLOps processes:
Data → Training → Evaluation → Deployment → Monitoring
• Develop model training pipelines, including experiment management and comparative evaluations
• Implement and manage Model Registry, versioning, and artifact management (MLflow / DVC)
• Deploy models to production environments (batch and real-time inference)
• Work with distributed infrastructure (Ray / Kubernetes) for large-scale training and inference
• Monitor model performance, drift, and data quality, driving continuous improvement
• Automate retraining and continuous evaluation workflows
• Collaborate closely with Researchers, Data Engineers, and DevOps teams to enable seamless transition from research to production
• Optimize performance (CPU/GPU) and infrastructure costs
Must Have
• Proven experience as an MLOps Engineer, ML Engineer, or Data Engineer with strong ML orientation
• Strong Python skills and hands-on experience with ML/DL frameworks (PyTorch / TensorFlow)
• Experience building model training and inference pipelines
• Experience with experiment tracking and model management tools (MLflow / Weights & Biases)
• Experience with Docker and Kubernetes
• Experience working in cloud environments (AWS / GCP / Azure)
• Solid understanding of the model lifecycle and challenges of moving models to production
• Ability to work effectively in a multidisciplinary, fast-paced environment
Nice to Have
• Experience working with satellite data (EO / SAR)
• Familiarity with Ray, Airflow, or Prefect
• Experience with GPU workloads and model training optimization
• Knowledge of observability tools (Prometheus, Grafana, OpenTelemetry)
• Experience with data versioning (DVC) or feature stores
• Experience with real-time or mission-critical systems
Summary
This is a key role at the intersection of research and operational deployment, responsible for transforming models into reliable, scalable, and monitored production capabilities.
Ideal for someone looking to combine AI, data, and infrastructure while working on complex systems with real-world impact.
Questions about this role
- This listing did not state a salary. We only show pay when the employer publishes it.
- Proven experience as an MLOps Engineer, ML Engineer, or Data Engineer with strong ML orientation, Strong Python skills, hands-on experience with ML/DL frameworks (PyTorch / TensorFlow), Experience building model training and inference pipelines, Experience with experiment tracking and model management tools (MLflow / Weights & Biases)