Senior Computer Vision Engineer (Egocentric), Data Foundry

Posted 21 Days Ago
2 Locations
In-Office or Remote
Senior level
Logistics • Software
The Role
Lead design and delivery of an egocentric video data product: capture rigs, edge processing, dataset pipelines, VLM-assisted annotation, and production-ready perception models (detection, tracking, segmentation, depth/3D reconstruction, and 6DoF pose estimation). Operate and scale warehouse capture programs and own model training, evaluation, optimization, and deployment.
Summary Generated by Built In

Stord is The Consumer Experience Company, powering seamless checkout through delivery for today's leading brands. Stord is rapidly growing and is on track to double our revenue in the next 18 months. To meet and exceed this target, Stord is strategically scaling teams across the entire company, and seeking energetic experts to help us achieve our mission.

By combining comprehensive commerce-enablement technology with high-volume fulfillment services, Stord provides brands a platform to compete with retail giants. Stord manages over $10 billion of commerce annually through its fulfillment, warehousing, transportation, and operator-built software suite including OMS, Pre- and Post-Purchase, and WMS platforms. Stord is leveling the playing field for all brands to deliver the best consumer experience at scale.

With Stord, brands can increase cart conversion, improve unit economics, and drive sustained customer loyalty. Stord’s end-to-end commerce solutions combine best-in-class omnichannel fulfillment and shipping with leading technology to ensure fast shipping, reliable delivery promises, easy access to more channels, and improved margins on every order.

Hundreds of leading DTC and B2B companies like AG1, True Classic, Native, Seed Health, quip, goodr, Sundays for Dogs, and more trust Stord to deliver industry-leading consumer experiences on every order. Stord is headquartered in Atlanta with facilities across the United States, Canada, and Europe. Stord is backed by top-tier investors including Kleiner Perkins, Franklin Templeton, Founders Fund, Strike Capital, Baillie Gifford, and Salesforce Ventures.

Stord operates one of the largest independent e-commerce fulfillment networks in the U.S., with 20+ fulfillment centers, 4,000+ warehouse associates, and nearly 100 million packages shipped annually.
We’re building a new business line that transforms this real-world operational infrastructure into high-value training data for the next generation of physical AI.
We’re looking for an experienced Computer Vision Engineer and technologist to help build and scale this business from the ground up. You’ll work at the intersection of computer vision, robotics, data, and warehouse operations, turning real-world environments and workflows into high-quality datasets that enable smarter, more capable AI systems.

Why This Role:

This is an opportunity to build a new physical AI data business from the ground up, with access to an operating environment that would be difficult to replicate anywhere else.

You’ll have:

  • A structural data advantage. Direct access to real-world warehouse environments, workflows, and human activity at significant scale.

  • A massive and rapidly growing market. Build data products for companies developing the next generation of robotics and physical AI.

  • True 0→1 ownership. Shape the product, technology, team, and operating model from the beginning.

  • Direct partnership with the CTO & Co-Founder. Work closely with company leadership to define the technical and commercial direction of the business.

What You'll Do:

You will own the early egocentric video and perception stack—from data collection and camera rigs through vision models, processing pipelines, and dataset delivery. This is a hands-on builder-operator role: you’ll define what needs to be built, build the critical pieces yourself, and work with a small team to operationalize and scale them.

You will:

  • Define and build the data product. Own the product across quality tiers—from RGB egocentric video to depth-enhanced and multimodal capture with hand pose, body pose, and annotations. Prioritize what gets built based on customer demand and hold a high bar for data quality.

  • Stand up the capture operation. Own camera and rig selection, hardware setup, enrollment, edge processing, data ingestion, and the pipelines that turn raw capture into production-ready datasets. Partner closely with warehouse operations, engineering, and customers to deliver on spec and on schedule.

  • Build the perception stack. Develop detection, tracking, segmentation, depth/3D reconstruction, 6DoF, and multi-view 3D hand/body pose estimation across egocentric and fixed-camera systems.

  • Automate labeling at scale. Build VLM-assisted and automated labeling workflows with human-in-the-loop QA, reducing the cost of annotation while maintaining rigorous quality standards.

  • Own the hardware–vision intersection. Drive camera calibration, epipolar and multi-view geometry, frame-accurate synchronization, and 3D pose triangulation across multi-camera and egocentric rigs.

  • Train and ship models. Design, fine-tune, optimize, and deploy computer vision and multimodal models against large, unstructured video datasets. Build reproducible systems that perform reliably in production—not models that live in a notebook.

Basic Qualifications:
  • 8+ years building and shipping production computer vision/perception systems, or an MS/PhD in computer vision, ML, or robotics with 6+ years of hands-on industry experience.

  • Experience building and scaling an egocentric perception or video data stack end to end, ideally within robotics, physical AI, or an AI data company.

  • Deep expertise in computer vision, including detection, tracking, segmentation, depth, 2D/3D pose estimation, and vision transformers.

  • Strong command of geometric computer vision, including camera calibration, multi-view geometry, synchronization, and 3D reconstruction.

  • Proven ownership of complex perception problems from data and model design through evaluation, optimization, and deployment, with measurable improvements in accuracy and reliability.

  • Experience working with large, unstructured video and multimodal datasets, with a disciplined approach to evaluation and quality measurement.

  • A track record of setting technical direction and raising the bar for other engineers.

  • Proven ability to take ambiguous 0→1 problems from concept to working system with limited resources and no established playbook.

  • Expert-level Python and strong software engineering fundamentals; C++ experience where performance demands it.

  • Most importantly, you’re a builder. You’re comfortable moving between hardware, data, models, infrastructure, and operations to make the system work.

Skills Required

  • Experience standing up and scaling an egocentric perception stack end-to-end (hardware, embedded perception, pipelines, delivery).
  • 8+ years building and shipping production computer-vision/perception systems (or MS/PhD plus 6+ years).
  • Deep expertise designing, training, and debugging CNNs and vision transformers.
  • Strong command of geometric computer vision: camera calibration, depth estimation, 2D/3D pose estimation.
  • End-to-end ownership of perception problems: data and model design through evaluation, optimization, and deployment with measurable outcomes.
  • Proven track record setting technical direction for a team or large workstream and raising engineering standards.
  • Ability to solve ambiguous 0->1 problems and deliver working systems with limited resources.
  • Experience with large unstructured video/multimodal datasets and rigorous evaluation instrumentation.
  • Expert Python and strong software-engineering fundamentals; C++ where performance demands it.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Atlanta, GA
222 Employees
Year Founded: 2015

What We Do

Stord is on a mission to migrate supply chains to the cloud—empowering brands to build sophisticated, agile, and integrated supply chains. Founded in 2015 and headquartered in the heart of Atlanta's vibrant tech community, Stord is pioneering the world's first Cloud Supply Chain. The Cloud Supply Chain is the convergence of the digital and physical elements of logistics. With Stord's Cloud Supply Chain, businesses can build, expand, and optimize their physical supply chain operations across freight, warehousing, and fulfillment, with the speed, flexibility, and ease of modern cloud software.

Similar Jobs

Snap Inc. Logo Snap Inc.

Senior Data Scientist

Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Remote or Hybrid
6 Locations
5000 Employees
162K-284K Annually

Snap Inc. Logo Snap Inc.

Data Scientist

Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Remote or Hybrid
6 Locations
5000 Employees
133K-235K Annually

Block Logo Block

Account Executive

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
In-Office or Remote
Seattle, WA, USA
12000 Employees
129K-233K Annually

Block Logo Block

Senior GRC Engineer

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
In-Office or Remote
8 Locations
12000 Employees
185K-327K Annually

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account