Senior Computer Vision Engineer|
Onsite (New York, NY) Full-time
Who we are:
Pakket builds remote control systems for warehouses. Specifically, we track customer’s forklifts, pallet jacks, people, and trucks across their facility IP cameras, project them onto a shared bird's-eye-view floor map, and stream live positions to a customer-facing backend.
Our stack is a real-time Python pipeline (YOLO/Ultralytics detection, multi-object tracking, homography-based BEV projection), deployed on an on-site GPU edge server.
The system runs live at customer sites today. We are looking for a hands-on builder who is interested in owning the perception and the core of the multi-camera tracking systems.
What you'll do:
Build and improve the real-time pipeline: multi-camera RTSP capture, GPU decode, detection, fusion, and tracking.
Improve tracking quality across cameras: association, track lifecycle, ID switches, occlusions, and track revival.
Fine-tune, evaluate, and deploy detection models end-to-end — dataset prep through TensorRT/ONNX deployment.
Optimize the compute budget: batching, quantization, latency/throughput on the edge GPU.
Build the evaluation tooling that proves improvements: offline replay, ground-truth scoring (MOTA/IDF1/HOTA-tier metrics), and controlled A/B ablations.
Own adjacent systems work: backend/API, deployment, observability, and internal admin tooling.
Diagnose hard production failures from logs, replays, and captured data.
Requirements:
4+ years of hands on experience building Computer Vision systems
Real-time multi-camera video: RTSP capture, FFmpeg/PyAV decode, GPU (NVDEC) decode, frame pacing, backpressure.
Object detection with YOLO/Ultralytics end-to-end: dataset prep, fine-tuning, and deployment; inference optimization (batching, ONNX/TensorRT, CUDA).
Camera geometry and calibration: intrinsics + lens distortion, homography estimation/application, projection to a shared ground plane; strong NumPy/OpenCV/SciPy.
Linux + remote GPU operations: SSH to an edge machine and a cloud VM, run systemd/Docker services.
Experience building a system that ran in production on real video, and can speak to common failure modes.
Nice-to-haves:
Warehouse/logistics or industrial CV domain experience
Tracking objects across multiple cameras
Small-model classification (MobileNet-style) and sequence smoothing (e.g. HMM).
Video encoding/distribution: NVENC, RTMP/HLS, MediaMTX.
Familiarity with a startup environment. We're a small team so everyone's job is broader than the title.
What success looks like in 90 days
You can run and score a full-day replay of forklift movement end-to-end, and you've shipped one measured improvement to calibration, tracking quality, or latency with before/after numbers.
You’ve gained familiarity with our detection, location and icon placement logic pipeline
Engineering peers and field staff trust your diagnoses and write-ups
Why Pakket?
You'll be a founding engineer: you will receive meaningful equity, and ownership/control of the core technical system from day one
Our team is small. While your core responsibility is computer vision, you will have the opportunity to work on any part of the stack that is meaningfully interesting to you (backend/API/calibration etc.)
The problem is real and the demand is large: warehouses are a $1T+ industry still running on spreadsheets and radios. Our customers run some of the largest facilities in the country and are demanding a solution to solve their lack of visibility
You'll see the direct impact of your work — your model will start being tested/run on a warehouse as soon as it's shipped, and you'll watch it get better every week.
Competitive salary in a funded and rapidly growing company, and the opportunity to help build the company from the ground up.
Skills Required
- 4+ years of hands-on experience building computer vision systems
- Experience with real-time multi-camera video, including RTSP capture, FFmpeg or PyAV decoding, GPU/NVDEC decoding, frame pacing, and backpressure
- End-to-end experience with YOLO/Ultralytics object detection, including dataset preparation, fine-tuning, deployment, batching, ONNX/TensorRT, and CUDA optimization
- Experience with camera geometry and calibration, including intrinsics, lens distortion, homography, and shared-ground-plane projection
- Strong experience with NumPy, OpenCV, and SciPy
- Linux and remote GPU operations, including SSH, systemd, and Docker services
- Experience building and operating a computer vision system in production on real video
- Warehouse, logistics, or industrial computer vision experience
- Experience tracking objects across multiple cameras
- Experience with small-model classification, such as MobileNet-style models, and sequence smoothing such as HMMs
- Experience with video encoding or distribution, including NVENC, RTMP, HLS, or MediaMTX
- Familiarity with a startup environment
What We Do
Visualize your Warehouse. With key metrics & analysis. Like never before. Pakket AI is a Warehouse Intelligence Platform that tracks forklifts, people, and trucks in real time.









