Skip to main content

AI Infrastructure Architect

NeuRealityCenter District, IsraelNot specifiedFull-timeSeniority: Lead

Posted 24 days ago · 58 applicants

Salary not listed for this role

Saving, applying or scoring takes a few seconds to set up your free account.

Willbi insight

The role in plain words

The Lead System Architect will define the software architecture and technical roadmap for NeuReality's next-generation AI inference platform, NR-NEXUS. Day-to-day responsibilities include writing system specifications, researching AI infrastructure trends, and collaborating with engineering teams to optimize performance and scalability for distributed AI workloads. The role also involves mentoring software engineers and defining performance goals.

Must-have
  • 7+ years of software engineering experience
  • 3+ years in software architecture or technical leadership
  • Deep understanding of Gen AI/LLM infrastructure and distributed workloads
  • Strong experience with Kubernetes-based platforms and cloud-native architecture
  • Experience designing management software or SaaS platforms for production systems
Nice-to-have
  • Deep understanding of GenAI / LLM inference infrastructure, including model serving, scaling, batching, latency, throughput, and resource utilization
  • Experience with production AI inference clusters using GPUs, AI accelerators, or other specialized compute infrastructure
  • Understanding of how distributed inference systems operate, including scheduling, load balancing, autoscaling, failover, and cluster-level observability
  • Experience with LLM serving frameworks such as vLLM, Triton Inference Server, TensorRT-LLM, or similar
  • Familiarity with GPU/accelerator orchestration, device plugins, resource scheduling, and cluster capacity planning

Extracted from the job description · kept up to date automatically

Who this suits

This role suits experienced software engineers with at least 7 years of experience, including 3 or more years in technical leadership, who have a deep understanding of Gen AI/LLM infrastructure and Kubernetes-based platforms. It is less ideal for those without experience in distributed systems, microservices, and cloud-native architectures.

Full job description

Original listing · kept for reference

NeuReality is seeking a Lead System Architect to join our system architecture team and help define NR-NEXUS, our next-generation AI inference platform.

Responsibilitie

• Lead the software architecture and technical roadmap for NeuReality’s NR-Nexus

• Write system specifications for NR-Nexus product

• Research AI infrastructure, SaaS platforms, model serving, and inference trend

• Work with engineering to translate technical capabilities into product value

• Work closely with engineering teams to optimize performance, scalability, and feature delivery

• Define performance goals and lead profiling, benchmarking, and optimization efforts for GenAI and distributed AI workloads

• Collaborate with customers, partners, and open-source communities to ensure ecosystem compatibility and adoption

• Mentor software engineers and provide technical leadership

Requirements

• 7+ years of software engineering experience, including 3+ years in software architecture or technical leadership.

• Deep understanding of Gen AI/LLM infrastructure and distributed workloads- Must

• Strong experience with Kubernetes-based platforms and cloud-native architecture.

• Experience designing management software or SaaS platforms for production systems.

• Strong background in distributed systems, microservices, APIs, and automation.

• Hands-on experience with observability stacks, monitoring, logging, alerting, and SLA/SLO tracking.

• Experience with CI/CD, deployment automation, upgrades, and rollback mechanisms.

• Good understanding of security, authentication, authorization, and integration with customer data center environments.

Nice to have

• Deep understanding of GenAI / LLM inference infrastructure, including model serving, scaling, batching, latency, throughput, and resource utilization.

• Experience with production AI inference clusters using GPUs, AI accelerators, or other specialized compute infrastructure.

• Understanding of how distributed inference systems operate, including scheduling, load balancing, autoscaling, failover, and cluster-level observability.

• Experience with LLM serving frameworks such as vLLM, Triton Inference Server, TensorRT-LLM, or similar.

• Familiarity with GPU/accelerator orchestration, device plugins, resource scheduling, and cluster capacity planning.

• Familiarity with GPU communication technologies such as GPUDirect RDMA, NCCL, NVLink, or UALink.

• Experience optimizing communication for distributed AI/ML workloads.

• Knowledge of Prometheus, Grafana, OpenTelemetry, Helm, Argo CD, Istio, KServe, Kubeflow, or similar tools.

• Experience deploying software in on-prem, edge, private cloud, or hybrid environments.

About NeuReality
Company profile · coming soon

Employee reviews · coming soonMore roles at NeuReality

Questions about this role

  • This listing did not state a salary. We only show pay when the employer publishes it.
NeuReality
Posted 24 days ago · 58 applicants
See how you match