- M.S. or Ph.D. in Data Science, Machine Learning, Computer Science, or equivalent practical experience in deep learning
- Strong hands-on experience training large-scale models using RL algorithms (e.g. GRPO, PPO) on open-weight architectures
- Solid understanding of LLM vulnerabilities, red-teaming methodologies, and defensive alignment against IPI attacks
- PyTorch
- Docker containerization
- Prior experience working with standard RL gym formats, such as Harbor
- Experience evaluating complex agentic workflows in tool-use or web-browser environments
- Familiarity with evaluating open-weight models similar to Llama or Mistral against adversarial workloads
חולץ מתיאור המשרה · מתעדכן אוטומטית
תיאור המשרה המלא
המשרה המקורית · נשמר לעיוןWe are seeking a Senior AI Researcher to lead post-training evaluation, red-teaming, and reinforcement learning (RL) gym audits on open-weight models. The ideal candidate will establish rigorous benchmarking methodologies, evaluate large language models (LLMs) against complex threats like Indirect Prompt Injections (IPI), and construct post-training evaluation pipelines that accurately measure realistic frontier-level security capabilities. Key Responsibilities • RL Post-Training & Benchmarking: Execute post-training runs (e.g. GRPO) using mainstream open-weight generalist models against security-focused RL environments, targeting threat vectors like Indirect Prompt Injection (IPI). Reward Diagnostics & Trace Analysis - Analyze live loss curves and rollout traces to identify reward hacking, lazy policy convergence, and flawed or over/under-specified verifiers. • Task & Environment Auditing: Review tasks and multi-turn environments (including tool use, web navigation, and computer use) for realism, threat model accuracy, data distribution, and dataset balance. • Performance Reporting (Gym Cards): Generate comprehensive evaluation cards detailing hill-climbing performance uplift across checkpoints, failure modes, tokens/turns per rollout, and task-level success rates. • Integration & Orchestration: Integrate dockerized environments (e.g., Harbor format) into internal training frameworks, optimizing reset/statefulness semantics, concurrency, and throughput ceilings.
Requirements: Required Qualifications • Technical Background: M.S. or Ph.D. in Data Science, Machine Learning, Computer Science, or equivalent practical experience in deep learning. • RL & Post-Training Expertise: Strong hands-on experience training large-scale models using RL algorithms (e.g. GRPO, PPO) on open-weight architectures. • AI Security Expertise: Solid understanding of LLM vulnerabilities, red-teaming methodologies, and defensive alignment against IPI attacks. • Infrastructure Skills: Proficiency in PyTorch, Docker containerization, and distributed training architectures. • Diagnostic Skills: Ability to analyze agent rollout traces, craft deterministic rubrics/verifiers, and debug complex reward shaping flaws. Preferred Qualifications • Prior experience working with standard RL gym formats, such as Harbor. • Experience evaluating complex agentic workflows in tool-use or web-browser environments. • Familiarity with evaluating open-weight models similar to Llama or Mistral against adversarial workloads.
שאלות על המשרה
- המשרה לא ציינה שכר. אנחנו מציגים שכר רק כשהמעסיק מפרסם אותו.
- היברידי
- M.S. or Ph.D. in Data Science, Machine Learning, Computer Science, or equivalent practical experience in deep learning, Strong hands-on experience training large-scale models using RL algorithms (e.g. GRPO, PPO) on open-weight architectures, Solid understanding of LLM vulnerabilities, red-teaming methodologies, and defensive alignment against IPI attacks, PyTorch, Docker containerization