Applicants must be authorized to work for ANY employer in the U.S. We are unable to sponsor or take over sponsorship of an employment Visa at this time.
Principal DevSecOps & AI Platform Engineer
Job Summary
We are seeking a highly experienced Senior/Principal Cloud & AI Platform Engineer to design, build, secure, and operate scalable cloud platforms and AI/ML solutions. This role combines software engineering, cloud infrastructure, DevSecOps, SRE, security engineering, and AI/ML expertise.
The ideal candidate will have strong hands-on experience with Python, TypeScript, AWS, Terraform, CI/CD, LLMs, and RAG architectures, along with a proven ability to improve reliability, observability, security, and operational processes. This individual will also provide technical leadership and collaborate across engineering, security, infrastructure, and product teams.
Key Responsibilities
- Design, develop, and maintain scalable cloud-native applications and platforms using Python and TypeScript.
- Architect and implement solutions within AWS, following cloud security, scalability, reliability, and performance best practices.
- Build and maintain infrastructure using Terraform and Infrastructure as Code (IaC) principles.
- Design, implement, and improve CI/CD pipelines to automate software delivery, testing, infrastructure deployment, and security checks.
- Apply DevSecOps practices by integrating security throughout the software development and deployment lifecycle.
- Partner with security teams to identify vulnerabilities, implement security controls, and improve cloud and application security.
- Apply SRE principles to improve system reliability, availability, scalability, and operational efficiency.
- Develop and support AI/ML solutions, including applications leveraging LLMs and Retrieval-Augmented Generation (RAG).
- Design AI-powered applications that integrate LLMs with enterprise data, APIs, and other systems.
- Implement observability solutions covering metrics, logs, traces, application performance, and infrastructure health.
- Participate in incident management, including troubleshooting production issues, root cause analysis, remediation, and prevention.
- Establish and improve monitoring, alerting, incident response, and operational procedures.
- Collaborate with engineering, DevOps, security, data, and business stakeholders to deliver enterprise solutions.
- Provide technical leadership, mentoring, and guidance to engineers and development teams.
- Identify opportunities to automate manual processes and improve engineering productivity.
- Establish engineering standards and best practices around cloud architecture, security, reliability, and AI/ML development.
Required Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- 7+ years of experience in software engineering, cloud engineering, DevOps, SRE, or a related technical discipline.
- Strong hands-on experience with Python.
- Experience developing applications or services using TypeScript.
- Strong experience with AWS cloud services and architecture.
- Hands-on experience with Terraform and Infrastructure as Code.
- Experience designing and maintaining CI/CD pipelines.
- Strong understanding of DevSecOps practices and cloud/application security.
- Experience with SRE principles, production operations, and reliability engineering.
- Experience with incident management, troubleshooting, and root cause analysis.
- Strong understanding of observability, including monitoring, logging, metrics, tracing, and alerting.
- Experience with AI/ML technologies, with practical experience working with LLMs and/or RAG architectures.
Kaleidoscope, an Infosys Company, is an equal opportunity employer, and all qualified applicants will receive consideration without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, protected veteran status, spouse of protected veteran, or disability.
Skills Required
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field
- 7+ years of experience in software engineering, cloud engineering, DevOps, SRE, or a related technical discipline
- Strong hands-on experience with Python
- Experience developing applications or services using TypeScript
- Strong experience with AWS cloud services and architecture
- Hands-on experience with Terraform and Infrastructure as Code
- Experience designing and maintaining CI/CD pipelines
- Strong understanding of DevSecOps practices and cloud/application security
- Experience with SRE principles, production operations, and reliability engineering
- Experience with incident management, troubleshooting, and root cause analysis
- Strong understanding of observability, including monitoring, logging, metrics, tracing, and alerting
- Experience with AI/ML technologies, including practical experience with LLMs and/or RAG architectures
What We Do
When clients come to us for product design and development, they get a full range of technical expertise and laboratory resources, but they also get a team that’s relentless when it comes to solving problems and creating designs that are the ideal combination of function and form. What does that mean? We build tools that save lives. For surgeons, it means they have tools that not only address but anticipate their needs, which helps improve their patients’ outcomes. We create products that save money and ensure safety. For people juggling family, work, and other demands, it means they have products that provide real help in managing their everyday lives. And for businesses, it means their bottom line is supported with durable tools that improve efficiency and protect their teams. From concept to production plan, or at any phase in between, we are a partner that can extend your product development know-how or provide additional mindpower to free up in-house teams.








