AI Performance Engineer JB-5254
פורסם 3 באוג׳ · 0 מועמדים
התפקיד במילים פשוטות
תפקיד זה מתמקד באופטימיזציה של ביצועי הסקת מסקנות (inference) עבור מודלי שפה גדולים (LLM) לכל אורך תשתית המערכת. העבודה היומיומית כוללת פיתוח ושיפור קרנלים ל-GPU, ניתוח נתוני פרופיילינג ושיפור יעילות הזיכרון, השהות והתפוקה ברמת הצומת הבודד וברמת האשכול המבוזר. בנוסף, התפקיד כולל שיתוף פעולה הדוק עם צוותי מחקר, תשתית ומוצר כדי להאיץ את השימוש במודלי AI בייצור.
- Strong experience in at least one of: performance engineering and low-level software optimization, GPU kernel development and optimization, or AI/ML systems optimization (especially inference/distributed serving)
- Strong programming skills in C++, CUDA, Python, or similar systems-oriented languages
- Experience profiling and optimizing software for latency, throughput, memory usage, or hardware utilization
- Familiarity with GPU architectures and performance characteristics
- Experience with one or more of CUDA, Triton, CUTLASS, ROCm, NCCL, TensorRT, XLA, TVM, vLLM, TensorRT-LLM, PyTorch, or similar technologies
- Experience working with large-scale distributed systems, high-performance computing, networking, or cluster scheduling
- Experience optimizing transformer-based models or production LLM inference systems
- Experience with multi-GPU or multi-node inference
- Familiarity with compiler-level optimization, graph optimization, or model runtime internals
- Experience with observability, benchmarking, and performance regression testing
חולץ מתיאור המשרה · מתעדכן אוטומטית
למי זה מתאים
התפקיד מתאים למהנדסי מערכות ותוכנה עם ניסיון באופטימיזציית תוכנה ברמה נמוכה, פיתוח קרנלים ל-GPU או אופטימיזציית מערכות AI/ML, בעלי שליטה ב-C++, CUDA או Python. הוא פחות מתאים למי שחסר רקע טכני בהנדסת ביצועים, ארכיטקטורת GPU או תשתיות מחשוב.
תיאור המשרה המלא
המשרה המקורית · נשמר לעיון• We are looking for an AI Performance Engineer to help optimize large language model inference across the full systems stack. You will work on improving the performance, efficiency, scalability, and reliability of LLM serving infrastructure, from low-level GPU kernel optimization to single-node runtime performance and distributed cluster-wide inference optimization.
• This role sits at the intersection of systems engineering, GPU computing, and AI infrastructure. You will work closely with researchers, ML engineers, infrastructure engineers, and product teams to make state-of-the-art AI models faster, more efficient, and easier to serve at scale.
• What You’ll Do
• Optimize LLM inference performance across the full stack, including kernels, runtimes, model execution, networking, scheduling, and distributed serving.
• Analyze and improve GPU utilization, memory bandwidth, latency, throughput, and cost efficiency.
• Develop, tune, or integrate high-performance GPU kernels using technologies such as CUDA, Triton, CUTLASS, or similar frameworks.
• Improve single-node inference performance through runtime optimization, memory management, batching, quantization, parallelism, and profiling.
• Optimize distributed inference across clusters, including tensor parallelism, pipeline parallelism, expert parallelism, communication patterns, load balancing, and scheduling.
• Build tools, benchmarks, and profiling workflows to identify bottlenecks and measure performance improvements.
• Collaborate with ML, infrastructure, and product teams to translate performance improvements into production impact.
• Stay current with advances in LLM serving, GPU architectures, compiler/runtime systems, and AI infrastructure.
• What We’re Looking For We are looking for candidates with strong experience in at least one of the following areas:
• Performance engineering and low-level software optimization.
• GPU kernel development and optimization.
• AI/ML systems optimization, especially for inference or distributed training/serving.
• Ideal Qualifications
• Strong programming skills in C++, CUDA, Python, or similar systems-oriented languages.
• Experience profiling and optimizing software for latency, throughput, memory usage, or hardware utilization.
• Familiarity with GPU architectures and performance characteristics.
• Experience with one or more of CUDA, Triton, CUTLASS, ROCm, NCCL, TensorRT, XLA, TVM, vLLM, TensorRT-LLM, PyTorch, or similar technologies.
• Understanding of LLM inference techniques such as batching, KV cache management, quantization, speculative decoding, parallelism strategies, and distributed serving.
• Experience working with large-scale distributed systems, high-performance computing, networking, or cluster scheduling is a plus.
• Ability to reason from first principles, use profiling data effectively, and drive measurable performance improvements.
• Strong communication skills and ability to collaborate across research, engineering, and infrastructure teams.
• Nice to Have:
• Experience optimizing transformer-based models or production LLM inference systems.
• Experience with multi-GPU or multi-node inference.
• Familiarity with compiler-level optimization, graph optimization, or model runtime internals.
• Experience with observability, benchmarking, and performance regression testing.
• Contributions to open-source AI infrastructure, GPU computing, or systems performance projects.
• Why Join Us:
• You will work on some of the most important performance challenges in modern AI infrastructure. Your work will directly improve the speed, efficiency, and scalability of LLM systems, enabling better user experiences and more cost-effective AI deployment at scale.
שאלות על המשרה
- המשרה לא ציינה שכר. אנחנו מציגים שכר רק כשהמעסיק מפרסם אותו.
- Strong experience in at least one of: performance engineering and low-level software optimization, GPU kernel development and optimization, or AI/ML systems optimization (especially inference/distributed serving), Strong programming skills in C++, CUDA, Python, or similar systems-oriented languages, Experience profiling and optimizing software for latency, throughput, memory usage, or hardware utilization, Familiarity with GPU architectures and performance characteristics, Experience with one or more of CUDA, Triton, CUTLASS, ROCm, NCCL, TensorRT, XLA, TVM, vLLM, TensorRT-LLM, PyTorch, or similar technologies