Skip to main content

SRE Technical Lead, DevOps Group

BigPandaTel Aviv-Yafo, Tel Aviv District, IsraelHybridFull-timeSeniority: Lead

Posted Jun 15 · 0 applicants

Salary not listed for this role

Saving, applying or scoring takes a few seconds to set up your free account.

Willbi insight

The role in plain words

This role is a hands-on individual contributor position focused on ensuring the reliability, scalability, and performance of the BigPanda platform. You will define and track service goals, own production reliability including on-call rotations, and collaborate with engineering teams to embed reliability across the stack. Day-to-day tasks include driving reliability initiatives, implementing automation, and resolving incidents end-to-end.

Must-have
  • 5+ years of experience as an SRE or Platform Developer (or similar) in a high-scale production environment
  • AWS
  • Strong coding skills and a software engineering mindset
  • Experience in building or integrating AI-driven solutions (e.g., LLMs, agents, or AI-powered operational tooling)
  • Experience with infrastructure-as-code and modern container orchestration platforms
Nice-to-have
  • Business-level reliability experience

Extracted from the job description · kept up to date automatically

Who this suits

This role suits experienced SREs or Platform Developers with five or more years of experience in high-scale AWS environments who have strong coding skills and experience with AI-driven solutions. It is less ideal for those seeking a people-management position, as this is a hands-on technical lead role.

Full job description

Original listing · kept for reference

Location requirements:

This role requires working out of the Tel Aviv office three days per week.

About The Role

• As a SRE Technical Lead at BigPanda, you will play a critical role in ensuring the reliability, scalability, and performance of the platform that powers our customers’ operations. You’ll operate at the intersection of software engineering and production operations, taking full ownership of the systems you build and run.

• This is a hands-on individual contributor role - not a managerial position. You'll be deep in the technical work, driving impact through engineering excellence rather than people management.

• This role is not just about responding to incidents - it’s about fundamentally improving how our platform behaves under real-world conditions. You will drive reliability initiatives end-to-end: defining measurable service goals, shaping engineering priorities through error budgets, and implementing solutions that prevent issues before they occur.

• You’ll work closely with teams across the organization, embedding reliability and observability into every layer of the stack. At the same time, you’ll leverage automation, modern infrastructure practices, and emerging AI capabilities to continuously evolve how we operate and scale.

What You Will Do

• Develop deep product knowledge across our platform - understanding its internals, failure modes, and operational behavior well enough to own incident resolution end-to-end.

• Define and track SLAs/SLOs/SLIs across critical platform services, and use error budgets to drive engineering decisions.

• Own production reliability - including on-call rotations, incident response, and post-mortems - with a focus on minimizing MTTR and preventing recurrence through systemic fixes, not just firefighting.

• Work hand-in-hand with engineering teams across the stack - infrastructure, application, and business layers - to embed reliability requirements everywhere.

What Skills And Experience You’ll Bring To BigPanda

• 5+ years of experience as an SRE or Platform Developer (or similar) in a high-scale production environment, with hands-on ownership across the full stack - infrastructure and application layers.

• Experience in designing, building, and operating cloud-native systems on AWS.

• Strong coding skills and a software engineering mindset - you build your own tools rather than waiting for someone else to.

• An experience in building or integrating AI-driven solutions (e.g., LLMs, agents, or AI-powered operational tooling).

• A true owner - you take responsibility for systems end-to-end and proactively drive improvements without waiting for direction.

• Business-level reliability experience is a strong advantage.

• Experience with infrastructure-as-code and modern container orchestration platforms.

About Us

BigPanda delivers agentic automation for IT operations. We enable enterprise IT to keep its digital world running by transforming manual and reactive human processes into intelligent, autonomous systems that detect, respond, and prevent IT incidents at machine speed. That’s why the world’s most trusted brands rely on BigPanda to improve operational efficiency and deliver exceptional service reliability to their customers.

We have an awesome team of motivated, knowledgeable, fun-loving, and friendly Pandas. We provide comprehensive health coverage, parental leave, competitive cash and equity compensation, and a supportive, collaborative, and innovative environment to empower you to do the best work of your career.

Our Benefits

• Competitive equity

• Hybrid work schedule

• Company funded health insurance

• 6 weeks fully paid Parental Leave

• Critical Family Medical Leave

• Financial planning services

• Employee learning & development budget

• Values-based recognition (quarterly and annually)

• Social community & ERG programs

• FreeFit gym package

• Work-life harmony

• Dog friendly office

About BigPanda
Company profile · coming soon

Employee reviews · coming soonMore roles at BigPanda

Questions about this role

  • This listing did not state a salary. We only show pay when the employer publishes it.
BigPanda
Posted Jun 15 · 0 applicants
See how you match