Computer Vision Engineer
Posted 27 days ago · 0 applicants
Saving, applying or scoring takes a few seconds to set up your free account.
The role in plain words
- Production-grade computer vision: object detection and segmentation shipped in real systems (not notebooks)
- Video / multi-frame understanding: tracking, temporal consistency across frames
- A track record of taking models to production and owning them
- Comfort with dirty, real-world capture and genuine evaluation rigor — you measure accuracy and reason about failure modes
- High agency: you spot a perception bottleneck and build the fix without waiting for a ticket
- Modern multimodal models / VLMs for reasoning over perception output
- Exposure to 3D / depth / point-cloud data
Extracted from the job description · kept up to date automatically
Who this suits
Full job description
Original listing · kept for referenceFounding-level · in the code · owning perception end to end.
We send technicians into thousands of retail locations (7-Eleven, Kroger, the largest enterprise portfolios in the world) and they walk every store with a camera. Your job is to turn that raw video into structured truth. You'll build the perception engine that watches a walkthrough and answers what a site survey used to need a human for: what's in the store, what condition it's in, and what it costs to replace.
The challenge
• Inputs: raw on-site video from retail walkthroughs — handheld, real-world lighting, occlusion, motion blur.
• Goal: extract structured store attributes (fixtures, equipment, conditions, counts, dimensions) and auto-answer survey questions.
• The catch: an ingestion pipeline that works at scale — not for one client but for hundreds of enterprise clients, each with different site attributes and physical layouts.
What you'll do
• Contribute to the video-to-attributes pipeline end to end: ingestion, detection, segmentation, multi-frame tracking.
• Train and ship models that recognize retail fixtures, equipment, and site conditions from real footage.
• Fuse perception output with multimodal reasoning to populate our canonical attribute registry.
• Keep it fast and cheap at portfolio scale — this runs across thousands of stores, not a demo.
Must have
• Production-grade computer vision: object detection and segmentation shipped in real systems (not notebooks).
• Video / multi-frame understanding: tracking, temporal consistency across frames.
• A track record of taking models to production and owning them.
• Comfort with dirty, real-world capture and genuine evaluation rigor — you measure accuracy and reason about failure modes.
• High agency: you spot a perception bottleneck and build the fix without waiting for a ticket.
Nice to have (or fast to learn)
• Modern multimodal models / VLMs for reasoning over perception output.
• Exposure to 3D / depth / point-cloud data.
• Retail or built-environment domain experience.
Do not apply if you need a ticket queue or detailed specs handed to you. You'll navigate ambiguity and build.
Questions about this role
- This listing did not state a salary. We only show pay when the employer publishes it.
- Production-grade computer vision: object detection and segmentation shipped in real systems (not notebooks), Video / multi-frame understanding: tracking, temporal consistency across frames, A track record of taking models to production and owning them, Comfort with dirty, real-world capture and genuine evaluation rigor — you measure accuracy and reason about failure modes, High agency: you spot a perception bottleneck and build the fix without waiting for a ticket