Senior Software Engineer, DevOps

Posted 25 Days Ago
Be an Early Applicant
Ann Arbor, MI, USA
In-Office
160K-190K Annually
Senior level
Artificial Intelligence • Software • Energy • Utilities
The Role
Lead design, build, and operation of off-device platform infrastructure across cloud and on-premise. Manage Kubernetes container deployments, AWS architecture, Terraform IaC, observability, CI/CD, incident response, capacity planning, and SOC2 compliance. Provide technical leadership, mentor engineers, and drive automation and scalability for ML and data workloads.
Summary Generated by Built In
Utilidata is a fast-growing NVIDIA-backed AI company enabling AI data centers to dynamically orchestrate power and unlock more compute capacity from existing energy infrastructure. For over a decade, we have applied AI to the electric grid — bringing real-time visibility and power-flow control to complex energy infrastructure. Our Karman platform, built on a custom NVIDIA module, brings that same capability to AI data centers, giving operators a way to better use the power already available to them.
We are seeking a Senior Software Engineer, DevOps to help design, build, and operate Utilidata's off-device platform that ingests, processes, and serves data flowing from edge AI devices. The role will build and maintain infrastructure across on-premises and cloud environments – bridging edge deployments with cloud-based data processing to support analytics, operations, and ML workloads at scale. This is a hands-on development role with deep technical ownership and cross-team visibility. This engineer will build and maintain the systems that keep our platform running, help shape infrastructure and deployment best practices, and support less experienced engineers. This engineer will partner closely with on-device and ML teams to ensure our off-device platform is resilient, well-instrumented, and ready to scale.
Responsibilities
  • Support the deployment and management of containerized applications using Kubernetes, ensuring optimal performance and availability
  • Contribute to strategic planning on how infrastructure solutions evolve to match Data Center partner requirements
  • Design, implement, and maintain scalable and reliable systems on AWS and/or on-premise
  • Utilize Terraform for infrastructure as code to automate the provisioning and management of cloud resources
  • Monitor system performance and uptime, ensuring systems meet established service level objectives (SLOs)
  • Support SOC2 security compliance requirements for data handling
  • Guide team members in DevOps practices, promoting a culture of reliability and excellence
  • Advocate for automation of operational tasks to enhance efficiency and reduce manual intervention
  • Collaborate with cross-functional teams to build and maintain CI/CD pipelines
  • Troubleshoot and resolve complex production issues, conducting root cause analysis and implementing corrective actions
  • Participate in on-call rotations and incident response teams
  • Assist in capacity planning, performance tuning, and technical decision-making
  • Drive continuous improvement initiatives for processes and infrastructure
Minimum Qualifications
  • 8+ years of development experience including experience in platform engineering, SRE, or distributed systems, with demonstrated senior-level impact
  • Experience designing and operating infrastructure across on-premises and cloud environments
  • Strong proficiency in container orchestration, particularly Kubernetes
  • Strong proficiency with AWS services and architecture
  • Hands-on experience with Terraform for infrastructure automation
  • Familiarity with monitoring tools (Prometheus, Grafana, or similar) and observability best practices
  • Strong problem-solving skills and attention to detail
  • Strong communication and collaboration skills, with experience contributing to technical outcomes
  • Willingness to travel up to 20% of time
Enhanced Qualifications (Nice to Have)
  • Bachelor's degree in Computer Science, Engineering, or a related field
  • Experience supporting or enabling MLOps platforms, model deployment pipelines, or ML-adjacent infrastructure
  • AI workload scheduling using Kubernetes
  • Knowledge of Apache Spark for large-scale data processing
  • Knowledge of database technologies (SQL, NoSQL)
  • Understanding of networking concepts and security best practices

Salary Range: $160,000 to $190,000 base compensation depending on experience and stock options. Salary will be commensurate with an individual's skills, training, years of experience, and in line with internal compensation bands. 
Location: This position is based at our company headquarters in Ann Arbor, Michigan, with flexibility for occasional remote work.
Our Commitments:
Utilidata values the diversity of our team. We provide equal employment opportunities without regard to race, color, religion, creed, sex, gender, sexual orientation, gender identity or expression, national origin, age, physical disability, mental disability, medical condition, pregnancy or childbirth, sexual orientation, genetics, genetic information, marital status, or status as a covered veteran or any other basis protected by applicable federal, state and local laws.
We are committed to:
  • Creating a diverse and inclusive workplace that is welcoming, supportive, affirming and respectful
  • Empowering employees to solve problems and work together to make a difference
  • Providing mentorship and growth opportunities as part of a collaborative team
  • A flexible work environment with flexible paid time off
  • Competitive compensation and benefits, including health, dental, vision, and employer-match 401k

Skills Required

  • 8+ years development experience in platform engineering, SRE, or distributed systems with senior/principal impact
  • Designing and operating infrastructure across on-premises and cloud environments
  • Strong proficiency in Kubernetes (container orchestration)
  • Strong proficiency with AWS services and architecture
  • Hands-on experience with Terraform for infrastructure automation
  • Familiarity with monitoring and observability tools (Prometheus, Grafana, or similar)
  • Excellent problem-solving, leadership, attention to detail, and ability to drive technical outcomes
  • Strong communication and collaboration skills
  • Willingness to travel up to 20%
  • Bachelor's degree in Computer Science, Engineering, or related field
  • Experience supporting or enabling MLOps platforms or ML-adjacent infrastructure
  • AI workload scheduling using Kubernetes
  • Knowledge of Apache Spark for large-scale data processing
  • Knowledge of database technologies (SQL, NoSQL)
  • Understanding of networking concepts and security best practices
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Providence, RI
77 Employees
Year Founded: 2012

What We Do

Utilidata is an AI-powered technology company that is working with NVIDIA to create the next generation of AI-embedded infrastructure, starting with the electric grid. Karman, our distributed AI platform, operates on our custom NVIDIA module, makes data available for accelerated computing at the edge, and trains AI models locally. Karman is embedded in grid devices - starting with smart meters - to transform the way utility companies operate. As the electric grid becomes more complex with the rapid increase of electric vehicles, distributed solar, batteries, heat pumps and extreme weather, utilities need real-time visibility of grid conditions and dynamic, software-defined infrastructure. Karman provides real-time visibility and AI at the grid edge so utilities can better utilize customer energy resources, reduce power outages, and enable quicker storm recovery. We are a mission-driven, collaborative, and adaptive team working to do what’s right, even when it’s hard. With backgrounds in electric engineering, power systems engineering, software engineering, data science, and energy policy, we bring a unique perspective on the solutions the energy industry needs. We are committed to ensuring a diverse, inclusive, and flexible workplace where employees are provided mentorship and growth opportunities and are empowered to solve problems as part of a collaborative team.

Similar Jobs

General Motors Logo General Motors

Senior Software Engineer

Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
Hybrid
2 Locations
165000 Employees
In-Office
2 Locations
175633 Employees
97K-193K Annually

Scribe Logo Scribe

Senior Software Engineer

Artificial Intelligence • Productivity • Software
In-Office or Remote
2 Locations
119 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account