דלג לתוכן הראשי

DevOps Production Engineer

Rhino Federated Computingתל אביב-יפו, מחוז תל אביב, ישראלהיברידיFull-timeדרגה: דרגת ביניים

פורסם לפני 11 ימים · 96 מועמדים

שכר לא צוין במשרה זו

שמירה, הגשה או בדיקת התאמה — כמה שניות להקמת חשבון חינם.

תובנת Willbi

התפקיד במילים פשוטות

מהנדס ה-Production יהיה אחראי על האמינות, הביצועים והתפעול השוטף של פלטפורמת המחשוב המבוזרת של החברה. התפקיד כולל תחזוקה ושיפור של סביבות ייצור בענן ומאחורי חומות אש של לקוחות, ניטור מערכות, פתרון תקלות וניהול אירועים בזמן אמת. בנוסף, התפקיד כולל עבודה צמודה עם צוותי פיתוח ופיתוח כלים לאוטומציה של שדרוגים והפצות.

חובה
  • 3–5 years of experience in Production Engineering, DevOps, SRE, or similar operational roles
יתרון
  • Familiarity with GitOps concepts and tools (ArgoCD)
  • Experience supporting AI/ML products, distributed systems, or customer-facing production systems

חולץ מתיאור המשרה · מתעדכן אוטומטית

למי זה מתאים

התפקיד מתאים לאנשי DevOps/SRE עם 3-5 שנות ניסיון בסביבות ענן, Kubernetes, Linux ופייתון, בעלי יכולות פתרון תקלות מורכבות. הוא פחות מתאים למי שמחפש תפקיד פיתוח טהור ללא היבטים תפעוליים וניהול אירועי ייצור.

תיאור המשרה המלא

המשרה המקורית · נשמר לעיון

About Rhino

Rhino Federated Computing Rhino solves one of the biggest challenges in AI: seamlessly connecting siloed data, apps, and models through federated computing. Our platform enables our customers to share business insight and processing without compromising intellectual property, data privacy, or data security. Our Rhino Federated Computing Platform offers flexible architecture across multi-cloud and on-prem hardware, end-to-end data management controls and workflows, and privacy enhancing technologies. Rhino is trusted by over 60 leading consortia and enterprises worldwide, including 14 of 20 of Newsweek’s ‘Best Smart Hospitals’ and top 20 global biopharma companies.

The company is headquartered in Boston, with an R&D center in Tel Aviv.

About The Role

The Production Engineer will play a key role in ensuring the reliability, performance, and operational excellence of Rhino’s Federated Computing Platform (Rhino FCP). This distributed infrastructure supports cutting-edge AI/ML research and development across highly regulated industries, including healthcare, finance, and life sciences, by enabling secure, privacy-preserving data collaboration worldwide.

You will be responsible for maintaining and improving production environments deployed both in cloud environments and behind customer firewalls. You will work closely with Platform, Backend, and Product teams to ensure systems remain stable, observable, and highly available.

This role focuses heavily on operational ownership, production monitoring, troubleshooting complex environments, incident response, and improving deployment reliability and operational tooling. It is ideal for someone who enjoys solving production challenges, improving system reliability, and building operational excellence in fast-moving environments.

Key Responsibilities

• Production Operations & Reliability: Maintain and support production environments across customer deployments and centralized cloud services, ensuring high availability and operational stability.

• Monitoring and Observability: Develop, improve, and maintain monitoring, alerting, and logging systems to proactively identify issues and improve visibility across distributed systems.

• Incident Response and Troubleshooting: Investigate, troubleshoot, and resolve complex infrastructure and application issues across cloud and on-premises environments, participating in incident management and root cause analysis.

• Deployment Management: Manage and support production deployments, upgrades, and maintenance activities across geographically distributed customer environments.

• Operational Excellence: Identify operational bottlenecks and continuously improve reliability, scalability, automation, and support processes.

• Collaboration Across Teams: Work closely with Backend, DevOps, and Product Engineering teams to support new features, improve operational readiness, and ensure smooth production adoption.

• Automation and Tooling: Contribute to internal tooling and automation efforts that reduce manual operational work and improve deployment and support efficiency.

Preferred Skills

Candidates should have 3–5 years of professional experience with a mix of the experiences described below:

• 3–5 years of experience in Production Engineering, DevOps, SRE, or similar operational roles

• Experience working with cloud environments (AWS and/or GCP preferred)

• Experience operating Kubernetes-based environments

• Strong Linux administration and troubleshooting skills

• Experience with Docker and containerized workloads

• Experience with Python and scripting for automation

• Experience with Infrastructure-as-Code and configuration management tools (Terraform, Ansible, or similar)

• Experience implementing and maintaining monitoring and observability systems (Prometheus/VictoriaMetrics, Grafana, alerting systems, logging pipelines, etc.)

• Experience working with CI/CD tools such as GitHub Actions

• Experience troubleshooting networking-related issues in distributed environments

• Familiarity with GitOps concepts and tools (ArgoCD is an advantage)

• Strong debugging and problem-solving abilities

• Comfortable working in dynamic startup environments

Bonus Skills

• Experience supporting AI/ML products or platforms

• Experience operating distributed systems

• Experience supporting customer-facing production systems

• Experience working with security-focused environments or privacy-sensitive workloads

• Experience with VPN technologies, mTLS, ingress systems, or service networking concepts

The role is open to candidates who are based in Israel (hybrid work environment).

אודות Rhino Federated Computing
פרופיל החברה · בקרוב

ביקורות עובדים · בקרובעוד משרות ב-Rhino Federated Computing

שאלות על המשרה

  • המשרה לא ציינה שכר. אנחנו מציגים שכר רק כשהמעסיק מפרסם אותו.
Rhino Federated Computing
פורסם לפני 11 ימים · 96 מועמדים
בדקו את ההתאמה