דלג לתוכן הראשי

Senior Platform Reliability Engineer

IgniteTechישראלמרחוקFull-timeדרגה: בכיר/ה

פורסם לפני 11 ימים · 0 מועמדים

שכר לא צוין במשרה זו

שמירה, הגשה או בדיקת התאמה — כמה שניות להקמת חשבון חינם.

תובנת Willbi

התפקיד במילים פשוטות

התפקיד כולל ניהול ותפעול של זמינות פלטפורמת ענן מבוססת בינה מלאכותית, תוך שימוש בסוכני AI לניהול תקלות ואוטומציה. העבודה היומיומית כוללת כוננויות, פתרון תקלות בזמן אמת, פיתוח תהליכי עבודה אוטונומיים וביצוע תחקירים מעמיקים למניעת תקלות חוזרות. בנוסף, התפקיד דורש הטמעת שינויי ייצור בטוחים ושיפור מתמיד של מערכות הניטור והאוטומציה.

חובה
  • 5+ years of hands-on production operations in SRE, Platform Engineering, DevOps, or Cloud Infrastructure at SaaS scale
  • AWS
  • Fluent, precise English
  • Committed to shift-based coverage and on-call rotations
  • OFAC-clear country of residence
יתרון
  • Grafana
  • Prometheus
  • Datadog
  • PagerDuty
  • OpsGenie

חולץ מתיאור המשרה · מתעדכן אוטומטית

למי זה מתאים

התפקיד מתאים לאנשי מקצוע בעלי 5 שנות ניסיון לפחות בתחומי SRE, DevOps או תשתיות ענן ב-AWS, עם ניסיון מעשי בניהול תקלות ייצור ושימוש בכלי בינה מלאכותית. הוא פחות יתאים למי שמחפש עבודה ללא תורנויות כוננות או למי שאין לו ניסיון מעשי מוכח בסביבות ענן מורכבות בקנה מידה גדול.

תיאור המשרה המלא

המשרה המקורית · נשמר לעיון

Senior Platform Reliability & AI Operations Engineer

Most reliability teams react to incidents. We engineer them out of existence.

We operate an AI-native community and social engagement platform that underpins customer relationships for Fortune 100 brands. When our platform goes down, it isn't a blip on an internal dashboard — it's a reputational crisis for some of the most recognized companies on earth. That context shapes everything about how we work, what we build, and who we hire.

We're looking for a senior reliability engineer who can carry production on their shoulders and build the autonomous systems that progressively carry it for them. You'll join a team where AI agents are first-class operational teammates — triaging alerts, validating changes, drafting root-cause analyses, and applying remediations within defined guardrails. Your mission is to make that surface area grow every single week.

Your Day-to-Day

• Carry the pager and own the outcome. You're the first responder on your shift window. When production degrades, you command the incident — diagnose, mitigate, escalate when blast radius demands it, and restore service. You treat every customer-impacting minute as personal accountability.

• Engineer autonomous operational workflows. Build, deploy, and refine the AI agents that handle pre-triage, change-gate validation, auto-healing, RCA drafting, and preventive-fix tracking. The agents are the product; your operational expertise is the training data.

• Ship safe production changes. Every deploy, config update, and cost-optimization action flows through quality gates with a validated rollback plan. You abort without hesitation the moment telemetry deviates from the expected path.

• Investigate to true root cause — then close the loop. Separate symptom from cause with disciplined analysis, then go further: identify the systemic prevention, build it, and track it to production. Unshipped RCA action items are unfinished work.

• Generalize every manual intervention. A one-off fix restores service; encoding it into an agent, runbook, or guardrail prevents recurrence. You're measured by how much the autonomous layer can handle — not by how many tickets pass through your hands.

• Multiply team knowledge. Encode procedures, context, and decision logic so agents can retrieve it and the next responder never starts from scratch. In a distributed, async organization, undocumented expertise doesn't count.

Who You Are

• 5+ years of hands-on production operations in SRE, Platform Engineering, DevOps, or Cloud Infrastructure at SaaS scale — with real first-responder incident experience, not adjacent project work.

• Battle-tested on AWS — multi-AZ, multi-account environments, infrastructure-as-code, production incident management, and change control with gates and rollbacks. You've weathered significant outages and carry the operational intuition that only comes from living through them.

• Radically self-directed. You run your shift like a founder runs a company. You identify gaps, prioritize ruthlessly, ship solutions, and raise the bar — without waiting for direction. If a standard is wrong, you challenge it openly; you never quietly ignore it.

• AI-native in practice, not in theory. You routinely delegate substantive operational work to agents, critically evaluate their output, and iterate on the underlying capabilities when they fall short. Experienced with agentic tooling — Claude Code, Codex, Warp, custom agent frameworks — and compulsively curious about new models and techniques as they emerge.

• AWS Solutions Architect – Associate or higher — or a production track record that renders the certification a formality.

• Fluent, precise English — in incident-bridge communication and in long-form writing alike.

• Committed to shift-based coverage. On-call rotations and your designated shift window are foundational to the role, with the time-zone overlap your window requires.

• OFAC-clear country of residence.

Bonus Points

• Original contributions to the agentic operations or AIOps space — open-source tooling, technical writing, conference talks, or shipped internal platforms.

• Production experience with multi-tenant B2B SaaS — community platforms, social tools, customer-experience products, or observability systems.

• Working knowledge of Grafana, Prometheus, Datadog, PagerDuty, or OpsGenie. Azure exposure alongside your AWS depth.

• Evidence of deep, sustained obsession with a hard problem — professional or personal. Depth of curiosity matters more than breadth of résumé.

What You'll Walk Away With

You won't just read about the future of AI-driven reliability — you'll be one of the engineers who built it, on a platform that Fortune 100 brands depend on daily. The agent-ops skills, incident patterns, and architectural instincts you develop here are things the broader industry is still struggling to define, let alone hire for.

How We Work

• Enterprise clients, startup velocity. Fortune 100 contractual stakes with a small-team cadence — weekly delivery cycles, fast decision-making, and a playbook that evolves as the field does.

• Uncapped tooling and compute. The agent harness is the product. If the right answer is a bigger model, more infrastructure, or a tool we haven't adopted yet, we invest.

• Fully remote. Global team.

אודות IgniteTech
פרופיל החברה · בקרוב

ביקורות עובדים · בקרובעוד משרות ב-IgniteTech

שאלות על המשרה

  • המשרה לא ציינה שכר. אנחנו מציגים שכר רק כשהמעסיק מפרסם אותו.
IgniteTech
פורסם לפני 11 ימים · 0 מועמדים
בדקו את ההתאמה