About this role
About Ironsite
Ironsite is building the intelligence layer for the physical world. We design our own wearable hardware, deploy it alongside craft workers, and transform a shift's footage into a next-morning report. Our internal team and purpose-built models label the data overnight and deliver actionable insights to superintendents by 5 AM.
We are accelerating the speed, efficiency, and predictability of construction, especially for complex, mission-critical infrastructure projects, including data centers, LNG facilities, sports stadiums, hospitals, and other large-scale developments, by training AI models on egocentric construction footage and labor productivity data. We are built with a pro-worker philosophy at our core: we believe technology should empower the workforce, not replace it. We're working to give craft workers and project leaders better visibility into what's happening on-site, while creating a system where the reality of construction and the chaos of each day is finally available to the people running the project.
Ironsite is deployed across several of the largest active construction projects in the country. To date, we've captured more than 100,000 hours of construction footage across seven states, now process thousands of hours of site activity every day, and maintain a worker opt-out rate below two percent. This is enabled by a workforce-first architecture that anonymizes devices, captures no audio, and never releases raw video.
Ironsite is backed by leading investors (8VC, South Park Commons, Saga Ventures) and prominent operators across technology and construction, including Eric Schmidt, Jeff Dean, Jeff Rothschild, Mark Leslie, Scott Wu, Eric Glyman, Karim Atiyeh, Russell Kaplan, and others, alongside over a dozen construction industry operators who have joined us as partners in building this.
Longer term, we believe Ironsite is the foundation for what construction becomes in the next decade. We think the systems we're building are the operating system for how the physical world gets built, and will unlock a fundamentally different way of respect for our workforce. One where craft workers are more valued, more visible, and better paid for the skill they bring, and where the industry finally has the intelligence layer that makes autonomous construction possible. Both futures start with the same foundation.
The Role
As Post-Training Lead, you'll report directly to our Chief Science Officer and own our post-training effort end-to-end: training, benchmarking, and deploying state-of-the-art VLMs that can interpret the complexity of a real construction site, built on data no other lab has.
This is a hands-on technical lead role. You'll train models yourself, lead the researchers already working on SFT and RL, and help grow the team around this effort. Model quality is the single most important output of this role.
Ironsite operates one of the most distinctive research environments in AI today. Our dataset is proprietary, growing by thousands of hours per day, expert-labeled, and structured around a taxonomy built for a specific real-world domain. Our compute footprint spans edge, on-prem, and cloud. Our models don't just ship to a benchmark; they ship to production, running on active jobsites within days of training. Very few post-training seats in the world offer that combination.
Open Problems You Could Own in Your First Year
- Post-training for expert-level perception. Using SFT and RL (GRPO and beyond) to push VLM labeling of fine-grained construction activity beyond expert human taggers, across every trade
- RLVR for agents that investigate the jobsite. Training agents with verifiable rewards to go beyond labeling and run their own research loop over a site's footage: form a hypothesis about what's slowing a crew down, dispatch a swarm to search thousands of hours of video for evidence, and return recommendations field leaders can act on the next morning
- Long-video understanding at industrial scale. Reasoning over multi-hour egocentric footage where the events that matter are sparse. This requires temporal grounding, long-context modeling, and memory beyond what current VLMs offer
- From understanding to prediction. Moving from models that describe what happened on a jobsite to models that forecast what happens next: how a crew's current sequence plays out over the coming days, where work will stall, and what changes would prevent it. The first step toward a true world model of construction
- Inference at fleet scale. In the next twelve months we'll collect and process millions of hours of video per month, eventually hundreds of thousands per day. Making frontier-quality inference cheap enough to run on all of it (distillation, quantization, model routing, hardware-aware optimization) is a partially unsolved problem you'd own
- Evaluation frameworks that predict reality. Building benchmarks that track real field performance as we expand to new jobsites, new trades, and even the same jobsite at a different stage of work. Standard academic benchmarks don't cut it for what we're solving
What You'll Do
- Architect and train novel VLMs. Design, train, and iterate on general-purpose vision-language models fine-tuned for spatial intelligence in construction environments. Set the bar for model quality across the research team
- Own the post-training roadmap. Take the lead on executing our research goals, from establishing baselines with state-of-the-art models to developing post-training recipes (SFT and RL), long-context architectures, and visual reasoning techniques
- Build on the Construction Intelligence Benchmark. Expand our existing benchmark suite across video question answering, temporal reasoning, activity recognition, and site-level analytical reasoning. Make it the reference point future construction AI research measures against
- Build scalable pipelines. Develop and own the model training and evaluation pipelines. Ensure we can rapidly experiment, measure performance, and deploy models into production
- Ship at production scale. Apply distillation, quantization, and model routing so state-of-the-art understanding runs affordably across millions of hours of monthly footage. Partner deeply with hardware and infrastructure teams so research decisions and deployment realities inform each other from day one
- Provide technical leadership. Set the technical direction for post-training, review designs and PRs, mentor researchers and interns, and set the standard for experimental rigor. This is technical leadership, not people management (though you'll be central to hiring the next generation of researchers)
Technical Challenges You'll Solve
- Training frontier models efficiently under real compute budgets while maximizing performance on the problems that matter for our customers
- Inventing new post-training objectives that capture construction-specific knowledge, temporal reasoning, and fine-grained perception at a level generic approaches can't reach
- Designing efficient attention mechanisms and architectural innovations for long-context understanding of construction workflows that span hours, days, and multi-project trajectories
- Building evaluation frameworks that measure real-world construction task performance beyond standard benchmarks, and defining what "good" means for a category the field hasn't yet formalized
- Balancing model capability with deployment constraints for edge, on-prem, and cloud inference across a fleet growing by orders of magnitude
- Working at the seams between research and production. The most interesting problems at Ironsite live where a training decision propagates all the way through to hardware constraints on a jobsite in Texas
What We're Looking For
Required
- Deep experience owning post-training (SFT and RL) for a large language or vision-language model that shipped to production. You've done this at least once, ideally more than once
- A track record of research contributions that meaningfully advanced the state of the art through papers, models, systems, or products others in the field have built on
- Deep expertise with modern deep learning frameworks (PyTorch, JAX, or similar) and strong proficiency in Python with solid software engineering fundamentals
- Deep experience working with and creating large-scale vision or language datasets
- Comfort operating at the frontier of what's known. You've led research on problems where the right answer wasn't in a paper yet, and you figured it out anyway
- Experience mentoring senior researchers and setting technical direction for a team
- A background in Computer Science, Machine Learning, AI, Robotics, or a related technical field, or the equivalent hands-on experience
Strongly preferred
- Deep expertise in fine-tuning and post-training large language or vision-language models (SFT, GRPO and other RL methods, parameter-efficient tuning such as LoRA, and novel post-training recipes)
- Hands-on experience with the hardest challenges of video data, including temporal reasoning, long-context modeling, memory, and efficient processing at scale
- Experience optimizing inference at industrial scale, including quantization, distillation, sparsity, and efficient serving on edge and cloud infrastructure
- Strong publication record at top-tier AI, ML, or CV conferences, or comparable evidence of research impact (open source, product influence, community leadership)
- Prior experience at a frontier lab, applied AI startup at scale, or research-heavy product company where you shipped post-training work to production
Nice to have
- Familiarity with MLOps tools for scalable model training and deployment
- Experience with multimodal models spanning vision, language, and additional modalities (audio, sensor, motion)
- Strong interest in vision-language models applied to real-world physical problems, and genuine curiosity about the day-to-day lives of construction workers
What Success Looks Like
- First 30 days: You know our data, our benchmarks, and our production models cold. You've identified the highest-leverage post-training directions for the next twelve months and made your case for what Ironsite should invest in and why
- First 3 months: You've led at least one major post-training initiative from problem definition through deployment. Your work has visibly moved the state of Ironsite's model performance on a problem that matters, and you're actively shaping the research roadmap alongside the Chief Science Officer
- First 6 months: You're one of the defining research voices at the company. You've mentored the rest of the research team, hired at least one senior researcher, and led the technical direction on a piece of Ironsite's post-training strategy that will run for years
Location, Compensation, & Perks
- San Francisco Bay Area (on-site)
- Base salary: $250k-$400k per year, commensurate with experience
- Significant early-stage equity
- Full benefits including health, dental, vision, and 401(k) with 6% match
- Access to dedicated GPU compute resources for research and experimentation
- Daily catered breakfast and lunch
- Office in San Francisco, next to Oracle Park and the Caltrain
Final compensation is determined by experience, location, and level.
Why You'll Love Working at Ironsite
- Foundational impact. Solve fundamental AI problems to transform one of the world's largest and least-digitized industries. Your models ship to real jobsites, not just papers
- Ownership and autonomy. As Post Training Lead, you'll have real authority over the post-training direction, technical decisions, and how the team operates. We value intellectual curiosity, first-principles thinking, and iterating quickly to turn ambitious ideas into reality
- Dream dataset. Exclusive access to a massive, proprietary, and continuously growing corpus of egocentric jobsite video from hundreds of devices deployed on active construction sites. A moat that enables frontier research no other lab can do
- World-class team. Lead alongside a small, elite team of researchers and engineers who have shipped cutting-edge AI products at companies like DeepMind, Etched, Meta, Apple, and NVIDIA
Tired of cold applications?
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
Know someone who'd be great for this?