DevOps Engineer - KubeOps Team
פורסם שלשום · 0 מועמדים
התפקיד במילים פשוטות
התפקיד כולל ניהול ותפעול סביבות Kubernetes בקנה מידה גדול (EKS) ללא הפרעה לאלפי השירותים הרצים עליהן. העבודה כוללת פיתוח תשתיות ענן ורכיבי ליבה של הקלסטרים כגון Karpenter, רשתות ו-autoscaling, לצד פיתוח כלי AI פנימיים ב-Go או Python ושמירה על יציבות המערכות.
- 2+ years in managing infrastructure, operating production systems at large scale
- 2+ years running production Kubernetes (EKS preferred) with strong understanding of core control-plane and cluster components, debugging, upgrading, tuning autoscaling and cluster networking
- 2+ years with a major cloud provider (AWS preferred) with solid grasp of compute, networking, and IAM
- 2+ years with Infrastructure-as-Code (Terraform, Pulumi, Crossplane, etc.), writing and maintaining reusable modules
- Observability at scale (Prometheus, Grafana, alert design)
- Cluster autoscaling in production
- Cost / capacity optimization at fleet scale (Karpenter consolidation, bin-packing, spot strategy)
- Linux internals and node-level debugging
- Track record of leading projects end-to-end: scoping, execution and delivery
חולץ מתיאור המשרה · מתעדכן אוטומטית
למי זה מתאים
התפקיד מתאים למהנדסי תשתיות עם שנתיים ומעלה של ניסיון בניהול מערכות פרודקשן גדולות, ידע מעמיק ב-Kubernetes, AWS ו-IaC, וניסיון בעבודה עם כלי AI. התפקיד פחות מתאים למי שמחפש עבודה מרחוק, שכן המשרה דורשת הגעה למשרד 5 ימים בשבוע.
תיאור המשרה המלא
המשרה המקורית · נשמר לעיוןJob Description
• Master the Cluster: Manage the lifecycle of our large-scale Kubernetes environments, without disrupting the thousands of services running on it.
• Build the Platform: Own and evolve the cluster building blocks — Kubernetes Architecture, Karpenter configurations, addons, networking, autoscaling. Where it makes sense, platformize them so other infra teams can consume them cleanly
• Pioneer AI-Driven Ops: Lead the charge in integrating AI into our ecosystem—building LLM-powered tools that accelerate investigation, planning, and coding for the whole team.
• Full-Stack Ownership: You’ll take ambiguous tasks and transform them into high-quality outcomes, owning every decision along the way.
Qualifications
• +2 years in managing infrastructure, operating production systems at large scale
• 2+ years running production Kubernetes — EKS preferred with a strong understanding of core control-plane and cluster components, with hands-on experience debugging, upgrading, and tuning autoscaling and cluster networking
• 2+ years with a major cloud provider (AWS preferred) — solid grasp of core compute, networking, and IAM.
• 2+ years with Infrastructure-as-Code, writing and maintaining reusable modules (Terraform, Pulumi, Crossplane, etc.)
• Familiarity with observability at scale — Prometheus, Grafana, alert design
• Experience building internal tooling in Go or Python
• Hands-on experience with AI-assisted development tools and agents (Claude Code, Codex, Cursor) for knowledge gathering, debugging, coding, planning, and design — with the ability to integrate these into concrete workflows that demonstrably accelerate your work.
Advantage:
• Cluster autoscaling in production
• Cost / capacity optimization at fleet scale — Karpenter consolidation, bin-packing, spot strategy
• Linux internals and node-level debugging — when a Kubernetes problem turns into a Linux problem, you can take it from there
• Track record of leading projects end-to-end: scoping, execution and delivery.
At Wix, we believe our best work happens together. Our work model is fully in person, with 5 days a week from our office. Flexibility remains a core value at Wix and special requests are handled thoughtfully at the team level.
Additional Information
KubeOps operates Wix's production Kubernetes platform at scale: a multi-DC EKS fleet with clusters running up to ~50,000 pods and ~2,000 nodes, hosting thousands of services across the company. We own the platform end-to-end — reliability, scale, upgrades, autoscaling and networking.
שאלות על המשרה
- המשרה לא ציינה שכר. אנחנו מציגים שכר רק כשהמעסיק מפרסם אותו.
- מהמשרד
- 2+ years in managing infrastructure, operating production systems at large scale, 2+ years running production Kubernetes (EKS preferred) with strong understanding of core control-plane and cluster components, debugging, upgrading, tuning autoscaling and cluster networking, 2+ years with a major cloud provider (AWS preferred) with solid grasp of compute, networking, and IAM, 2+ years with Infrastructure-as-Code (Terraform, Pulumi, Crossplane, etc.), writing and maintaining reusable modules, Observability at scale (Prometheus, Grafana, alert design)