The Role
Research Scientist developing vision-language models for construction-site intelligence. Responsibilities include training and evaluating deep learning models, building video and multimodal benchmarks, designing real-world performance metrics, and deploying optimized models to edge, on-premises, and cloud environments. The role involves end-to-end experimentation, long-context video reasoning, post-training, inference optimization, production feedback analysis, and collaboration with research, infrastructure, and hardware teams.
Summary Generated by Built In
About Ironsite
The Role
What You'll Do
Technical Challenges You'll Solve
What We're Looking For
What Success Looks Like
Location, Compensation, & Perks
Why You'll Love Working at Ironsite
Ironsite is building the intelligence layer for the physical world. We design our own wearable hardware, deploy it alongside craft workers, and transform a shift's footage into a next-morning report. Our internal team and purpose-built models label the data overnight and deliver actionable insights to superintendents by 5 AM.
We are accelerating the speed, efficiency, and predictability of construction, especially for complex, mission-critical infrastructure projects, including data centers, LNG facilities, sports stadiums, hospitals, and other large-scale developments, by training AI models on egocentric construction footage and labor productivity data. We are built with a pro-worker philosophy at our core: we believe technology should empower the workforce, not replace it. We're working to give craft workers and project leaders better visibility into what's happening on-site, while creating a system where the reality of construction and the chaos of each day is finally available to the people running the project.
Ironsite is deployed across several of the largest active construction projects in the country. To date, we've captured more than 100,000 hours of construction footage across seven states, now process thousands of hours of site activity every day, and maintain a worker opt-out rate below two percent. This is enabled by a workforce-first architecture that anonymizes devices, captures no audio, and never releases raw video.
Ironsite is backed by leading investors (8VC, South Park Commons, Saga Ventures) and prominent operators across technology and construction, including Eric Schmidt, Jeff Dean, Jeff Rothschild, Mark Leslie, Scott Wu, Eric Glyman, Karim Atiyeh, Russell Kaplan, and others, alongside over a dozen construction industry operators who have joined us as partners in building this.
Longer term, we believe Ironsite is the foundation for what construction becomes in the next decade. We think the systems we're building are the operating system for how the physical world gets built, and will unlock a fundamentally different way of respect for our workforce. One where craft workers are more valued, more visible, and better paid for the skill they bring, and where the industry finally has the intelligence layer that makes autonomous construction possible. Both futures start with the same foundation.
We're hiring a Research Scientist to help build the vision-language models that turn Ironsite's data into the intelligence layer we're building.
You'll work alongside our Chief Science Officer and the rest of the research team on the training, benchmarking, and deployment of state-of-the-art VLMs that can interpret the complexity of a real construction site. Your models will run on data no other lab has, and they'll ship to jobsites where they actually change how things get built.
Ironsite operates one of the most distinctive research environments in AI today. Our dataset is proprietary, growing by thousands of hours per day, expert-labeled, and structured around a taxonomy built for a specific real-world domain. Our compute footprint spans edge, on-prem, and cloud. And our models don't just ship to a benchmark, they ship to production, running on active jobsites within days of training. Very few research seats in the world offer that combination, and this is the earliest point at which you can join and shape the foundation of what we build.
What You'll Do
Train and iterate on Ironsite's core models.
- Design, train, and iterate on vision-language models fine-tuned for spatial intelligence in construction environments. Your work will directly determine how good our models get.
- Run experiments end to end, from data preparation through training, evaluation, and post-training. Move fast, measure honestly, and share what you learn with the team.
- Contribute to the frontier of what's possible on our data, whether that's establishing new baselines, improving post-training recipes, or exploring long-context architectures.
Build and improve our benchmarks.
- Contribute to our Construction Intelligence Benchmark suite across video question answering, temporal reasoning, activity recognition, and site-level analytical reasoning.
- Design evaluation metrics that measure real-world construction task performance, not just standard academic benchmarks.
- Own the evaluation loop for your own work so we always know what's actually improving.
Ship models to production.
- Take your best models from research code to production deployment, working closely with our infrastructure and hardware teams.
- Apply distillation, quantization, and model routing techniques so state-of-the-art understanding runs affordably across our growing fleet.
- Learn from what happens in the field. Field data and production feedback should shape your next experiment.
Grow with the team.
- Learn from senior researchers, contribute to the intellectual culture of the team, and start to develop your own point of view on what Ironsite's research should look like at scale.
- Read papers, share ideas, and help set the technical bar for the whole team.
- Training models efficiently under real compute budgets while maximizing performance on the problems that matter for our customers.
- Working with novel pre-training and post-training objectives that capture construction-specific knowledge, temporal reasoning, and fine-grained perception.
- Handling the challenges of long-context video data, including temporal reasoning, memory across multi-hour footage, and efficient processing.
- Designing evaluation metrics that predict real-world construction task performance.
- Balancing model capability with deployment constraints for edge, on-prem, and cloud inference.
What We're Looking For
Required
- 2-4 years of hands-on research experience designing and training deep learning models, particularly transformer-based architectures. Industry experience, top-tier PhD program, or equivalent.
- Deep expertise with modern deep learning frameworks (PyTorch, JAX, or similar) and strong proficiency in Python with solid software engineering fundamentals.
- Experience working with large-scale vision or language datasets.
- A track record of shipping meaningful research results, whether at a company, in a lab, or in publications.
- A background in Computer Science, Machine Learning, AI, Robotics, or a related field.
Strongly preferred
- Experience with vision-language models, video understanding, or multimodal architectures.
- Hands-on experience with post-training techniques for large language or vision-language models (SFT, RL methods, parameter-efficient tuning such as LoRA).
- Familiarity with the challenges of video data, including temporal reasoning and long-context modeling.
- Publications at top-tier AI, ML, or CV conferences.
Nice to have
- Experience optimizing inference at scale (quantization, distillation, sparsity).
- Familiarity with MLOps tools for model training and deployment.
- Interest in vision-language models applied to real-world physical problems, and genuine curiosity about the day-to-day lives of construction workers.
- First 30 days. You know our data, our benchmarks, and our production models cold. You've completed your first end-to-end experiment and shipped at least one meaningful improvement to a model in production.
- First 3 months. You've owned a research initiative from problem definition through deployment, and your work has visibly moved model performance on a problem that matters.
- First 6 months. You're a full contributor to the research team's roadmap and a trusted collaborator across the company. You've grown from executing on research problems to helping shape which ones we take on next.
- San Francisco Bay Area (on-site)
- Base salary: $175k-$275k per year, commensurate with experience
- Significant early-stage equity
- Full benefits including health, dental, vision, and 401(k) with 6% match
- Access to dedicated GPU compute resources for research and experimentation
- Daily catered breakfast and lunch
- Office in San Francisco, next to Oracle Park and the Caltrain
Final compensation is determined by experience, location, and level.
- Foundational impact. Solve fundamental AI problems to transform one of the world's largest and least-digitized industries. Your models ship to real jobsites, not just papers.
- Growth opportunity. We're at the stage where every researcher's work materially shapes the company. You'll get scope early, and the growth trajectory is defined by how much impact you're willing to have.
- Dream dataset. Exclusive access to a massive, proprietary, and continuously growing corpus of egocentric jobsite video from hundreds of devices deployed on active construction sites. A moat that enables frontier research.
- World-class team. Learn from and collaborate with a small, elite team of researchers and engineers who have shipped cutting-edge AI products at companies like DeepMind, Etched, Meta, Apple, and NVIDIA.
The base pay range for this role is $175,000 – $275,000 per year.
Skills Required
- 2-4 years of hands-on research experience designing and training deep learning models, particularly transformer-based architectures
- Industry experience, enrollment in a top-tier PhD program, or equivalent experience
- Deep expertise with modern deep learning frameworks such as PyTorch or JAX
- Strong proficiency in Python and solid software engineering fundamentals
- Experience working with large-scale vision or language datasets
- Track record of shipping meaningful research results in industry, a lab, or publications
- Background in Computer Science, Machine Learning, AI, Robotics, or a related field
- Experience with vision-language models, video understanding, or multimodal architectures
- Hands-on experience with post-training techniques including SFT, reinforcement learning methods, or parameter-efficient tuning such as LoRA
- Familiarity with video data, temporal reasoning, and long-context modeling
- Publications at top-tier AI, ML, or computer vision conferences
- Experience optimizing inference at scale using quantization, distillation, or sparsity
- Familiarity with MLOps tools for model training and deployment
- Interest in applying vision-language models to real-world physical problems
- Curiosity about the day-to-day lives of construction workers
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
Ironsite AI is a construction technology company that leverages wearable cameras and AI vision models to drive on-site productivity, safety, and training. By equipping workers with smart hard hats, the platform captures real-time data to analyze field activities, optimize labor allocation, and identify safety risks. Their mission is to modernize construction management by providing data-driven insights that help contractors reduce labor costs and deliver projects more efficiently.






