Job Description
- Provide embedded operations support, capability integration, and rapid prototyping for AI-enabled tools in support of our customer.
- Design, implement, and maintain robust infrastructure for enterprise AI applications in cloud environments (AWS, Microsoft Azure) and on-prem/air-gapped/classified environments.
- Develop and optimize engineering workflows and processes to support AI model and agent development, deployment, and maintenance.
- Architect and manage CI/CD pipelines for continuous integration and continuous delivery of AI models, agents, and applications.
- Implement and manage containerization and orchestration solutions using Docker and Kubernetes, and build and maintain control-plane guardrails — identity, policy enforcement, approvals, audit/logging, and observability.
- Ensure efficient AI model and agent lifecycle management, including versioning, monitoring, and scaling.
- Collaborate with AI/ML engineers and data scientists to streamline deployment processes and optimize resource utilization.
- Oversee system performance, security, and scalability of AI infrastructure, and support security accreditation (RMF/ATO) of AI systems.
- Continuously research and implement new DevOps tools and practices to enhance efficiency.
Our team is standing up an AI Enterprise capability — the central hub that mission teams turn to when they want to move an AI solution from prototype to a secure, governed, enterprise-scale service. As a Senior Platform / DevOps Engineer on this team, you will build and run the infrastructure, pipelines, and control-plane guardrails that let the command adopt AI safely and at scale, with significant, visible impact on the mission.
What You'll Be Doing:
Basic Qualifications
- Bachelor's degree in a relevant technical field with 10+ years of experience, or Master's degree in a relevant technical field with 8+ years of experience
- Advanced proficiency in DevOps principles and practices
- Demonstrated expertise in containerization using Docker and Kubernetes
- Proven experience in architecting and managing CI/CD pipelines
- Extensive experience with AI model lifecycle management and maintenance
- Active TS/SCI clearance w/ Poly
Preferred Skills
- Familiarity with cloud platforms (AWS, Microsoft Azure) for infrastructure deployment and management
- Familiarity with monitoring and logging tools (e.g., Prometheus, Grafana, ELK stack)
- Excellent communication and interpersonal skills, with the ability to effectively collaborate with cross-functional teams
- Experience with infrastructure as code (IaC) tools (e.g., Terraform, Ansible)
- Understanding of machine learning concepts and their implications for infrastructure; continuous learning mindset to stay abreast of cutting-edge DevOps and AI advancements
- Happy - Be Infectious. Happiness multiplies and creates a positive and connected environment where motivation and satisfaction have an outsized effect on everything we do.
- Helpful - Be Supportive. Being helpful is the foundation of teamwork, resulting in a supportive atmosphere where collaboration flourishes, and collective success is celebrated.
- Honest - Be Trustworthy. Honesty serves as our compass, ensuring transparent communication and ethical conduct, essential to who we are and the complex domains we support.
- Humble - Be Grounded. Success is not achieved alone, humility ensures a culture of mutual respect, encouraging open communication, and a willingness to learn from one another and take on any task.
- Hungry - Be Eager. Our hunger for excellence drives an insatiable appetite for innovation and continuous improvement, propelling us forward in the face of new and unprecedented challenges.
- Hustle - Be Driven. Hustle is reflected in our relentless work ethic, where we are each committed to going above and beyond to advance the mission and achieve success.
Skills Required
- Bachelor’s degree in a relevant technical field with 10+ years of experience, or a relevant master’s degree with 8+ years of experience
- Advanced proficiency in DevOps principles and practices
- Expertise with Docker and Kubernetes containerization
- Experience architecting and managing CI/CD pipelines
- Experience with AI model lifecycle management and maintenance
- Active TS/SCI clearance with polygraph
- Experience with AWS or Microsoft Azure
- Experience with monitoring and logging tools such as Prometheus, Grafana, or ELK Stack
- Experience with infrastructure as code tools such as Terraform or Ansible
- Understanding of machine learning concepts and infrastructure implications
- Excellent communication and cross-functional collaboration skills
What We Do
Agile Defense is a technology services company that provides advanced digital transformation, data analytics, and cybersecurity solutions to support critical national security and civilian government missions. With a global presence, the company focuses on delivering outcome-driven, AI-powered capabilities to solve complex mission challenges for federal and defense customers.









