Research Intern — Human Pose Understanding and Vision-Language Models
Posted 17 days ago · 0 applicants
Saving, applying or scoring takes a few seconds to set up your free account.
The role in plain words
This research intern role focuses on designing and implementing novel systems that integrate human pose representations with vision-language models. The selected candidate will collaborate with researchers and engineers to benchmark these methods and work toward a publication-ready contribution for a top-tier venue.
- Currently enrolled in a graduate program (M.Sc. or Ph.D.) in Computer Science, Electrical Engineering, or a related field
- Publications at top-tier venues (e.g., NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV, ACL, EMNLP or similar)
- Strong programming skills in Python and experience with deep learning frameworks (e.g., PyTorch)
- Solid foundation in computer vision, natural language processing, or multimodal learning
- Demonstrated expertise working with Vision-Language Models (VLMs) and/or Large Language Models (LLMs)
- Experience with human pose estimation, motion modeling, or related body-tracking tasks
- Familiarity with video understanding tasks and temporal modeling
- Familiarity with multimodal learning and benchmarks that combine language with visual or spatial data
- Experience with prompt engineering and optimization techniques
Extracted from the job description · kept up to date automatically
Who this suits
This role suits graduate students currently enrolled in an M.Sc. or Ph.D. program who have prior publications at top-tier computer vision or machine learning conferences and strong Python skills. It is less ideal for those without a solid foundation in computer vision, natural language processing, or multimodal learning.
Full job description
Original listing · kept for referenceSummary We are looking for a research intern to join us for a research project aimed at publication at a top-tier venue. The intern will design and develop novel systems that explore the interaction between human pose understanding and vision-language models (VLMs), advancing how these modalities can be combined to reason about human motion, activity, and embodied behavior across images and video.
Description Our group develops hand and body pose tracking algorithms for various apple devices and applications. One such example includes the hand tracking input for the Vision Pro.
Responsibilities
• Design and implement novel methods that integrate pose representations with vision-language models, targeting established academic benchmarks
• Collaborate with researchers and engineers on the team to produce a publication-ready contribution
• Benchmark against established evaluation suites and iterate toward state-of-the-art results
Minimum Qualifications
• Currently enrolled in a graduate program (M.Sc. or Ph.D.) in Computer Science, Electrical Engineering, or a related field
• Publications at top-tier venues (e.g., NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV, ACL, EMNLP or similar)
• Strong programming skills in Python and experience with deep learning frameworks (e.g., PyTorch)
• Solid foundation in computer vision, natural language processing, or multimodal learning
Preferred Qualifications
• Demonstrated expertise working with Vision-Language Models (VLMs) and/or Large Language Models (LLMs)
• Experience with human pose estimation, motion modeling, or related body-tracking tasks
• Familiarity with video understanding tasks and temporal modeling
• Familiarity with multimodal learning and benchmarks that combine language with visual or spatial data
• Experience with prompt engineering and optimization techniques
At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.
Learn about accessibility in Apple’s workplace
Role Number: 200671694-0865
Questions about this role
- This listing did not state a salary. We only show pay when the employer publishes it.
- Currently enrolled in a graduate program (M.Sc. or Ph.D.) in Computer Science, Electrical Engineering, or a related field, Publications at top-tier venues (e.g., NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV, ACL, EMNLP or similar), Strong programming skills in Python and experience with deep learning frameworks (e.g., PyTorch), Solid foundation in computer vision, natural language processing, or multimodal learning