MLOps Engineer

Posted 5 Days Ago
Be an Early Applicant
Arlington, VA, USA
In-Office
Senior level
Aerospace • Information Technology • Cybersecurity • Defense
The Role
Designs and maintains production machine learning infrastructure, including model deployment, monitoring, drift detection, centralized logging, CI/CD automation, security controls, and performance optimization. Builds scalable dashboards and pipelines using cloud, containerization, observability, and MLOps tools. Collaborates with data science and DevOps teams to ensure reliable, secure, and compliant AI systems.
Summary Generated by Built In
Overview/ Job Responsibilities

Job Summary

We are seeking a skilled MLOps Engineer to join our team and ensure the seamless deployment, monitoring, and optimization of AI models in production.

The MLOps Engineer will design, implement, and maintain end-to-end machine learning pipelines, focusing on automating model deployment, monitoring model health, detecting data drift, and managing AI-related logging. This role will involve building scalable infrastructure and dashboards for real-time and historical insights, ensuring models are secure, performant, and aligned with business needs.


Key Responsibilities

  • Model Deployment: Deploy and manage machine learning models in production using tools like MLflow, Kubeflow, or AWS SageMaker, ensuring scalability and low latency.
  • Monitoring and Observability: Build and maintain dashboards using Grafana, Prometheus, or Kibana to track real-time model health (e.g., accuracy, latency) and historical trends.
  • Data Drift Detection: Implement drift detection pipelines using tools like Evidently AI or Alibi Detect to identify shifts in data distributions and trigger alerts or retraining.
  • Logging and Tracing: Set up centralized logging with ELK Stack or OpenTelemetry to capture AI inference events, errors, and audit trails for debugging and compliance.
  • Pipeline Automation: Develop CI/CD pipelines with GitHub Actions or Jenkins to automate model updates, testing, and deployment.
  • Security and Compliance: Apply secure-by-design principles to protect data pipelines and models, using encryption, access controls, and compliance with regulations like GDPR or NIST AI RMF.
  • Collaboration: Work with data scientists, AI Integration Engineers, and DevOps teams to align model performance with business requirements and infrastructure capabilities.
  • Optimization: Optimize models for production (e.g., via quantization or pruning) and ensure efficient resource usage on cloud platforms like AWS, Azure, or Google Cloud.
  • Documentation: Maintain clear documentation of pipelines, dashboards, and monitoring processes for cross-team transparency. 
Minimum Qualifications

Qualifications

  • Education: Bachelor’s or Master’s degree in Computer Science, Data Science, Engineering, or a related field.
  • Experience:
    • 5+ years in MLOps, DevOps, or software engineering with a focus on AI/ML systems.
    • Proven experience deploying models in production using MLflow, Kubeflow, or cloud platforms (AWS SageMaker, Azure ML).
    • Hands-on experience with observability tools like Prometheus, Grafana, or Datadog for real-time monitoring.
  • Technical Skills:
    • Proficiency in Python and SQL; familiarity with JavaScript or Go is a plus.
    • Expertise in containerization (Docker, Kubernetes) and CI/CD tools (GitHub Actions, Jenkins).
    • Knowledge of time-series databases (e.g., InfluxDB, TimescaleDB) and logging frameworks (e.g., ELK Stack, OpenTelemetry).
    • Experience with drift detection tools (e.g., Evidently AI, Alibi Detect) and visualization libraries (e.g., Plotly, Seaborn).
  • AI-Specific Skills:
    • Understanding of model performance metrics (e.g., precision, recall, AUC) and drift detection methods (e.g., KS test, PSI).
    • Familiarity with AI vulnerabilities (e.g., data poisoning, adversarial attacks) and mitigation tools like Adversarial Robustness Toolbox (ART).
  • Soft Skills:
    • Strong problem-solving and debugging skills for resolving pipeline and monitoring issues.
    • Excellent collaboration and communication skills to work with cross-functional teams.
    • Attention to detail for ensuring accurate and secure dashboard reporting.
  • Must be eligible to obtain a Department of Homeland Security EOD clearance ( Requirements 1. US Citizenship, 2. Favorable Background Investigation) 
Desired Qualifications

Preferred Qualifications

  • Experience with LLM monitoring tools like LangSmith or Helicone for generative AI applications.
  • Knowledge of compliance frameworks (e.g., GDPR, HIPAA) for secure data handling.
  • Contributions to open-source MLOps projects or familiarity with X platform discussions on #MLOps or #AIOps.
About Us

Formed through the strategic union of Sev1Tech and ERT, Entarian is a premier provider of mission-critical engineering and technology solutions. Founded on a legacy of excellence dating back to 1993, Entarian is a product of an evolved and fully diversified engineering and federal technology leader. From deep space to defense and civilian missions, Entarian delivers secure, mission-aligned digital solutions that drive national resilience and operational effectiveness. We don't just support modernization; we define it.


Join the Mission and Start your Career Journey: Apply Directly via our Careers Portal  Connect, Referrals & Inquiries? Email the team: [email protected]


Entarian is an Equal Opportunity and Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, pregnancy, sexual orientation, gender identity, national origin, age, protected veteran status, or disability status.

Skills Required

  • Bachelor’s or Master’s degree in Computer Science, Data Science, Engineering, or a related field
  • 5+ years of experience in MLOps, DevOps, or software engineering focused on AI/ML systems
  • Production model deployment experience using MLflow, Kubeflow, AWS SageMaker, or Azure Machine Learning
  • Experience with observability tools such as Prometheus, Grafana, or Datadog
  • Proficiency in Python and SQL
  • Experience with Docker, Kubernetes, and CI/CD tools such as GitHub Actions or Jenkins
  • Knowledge of time-series databases and centralized logging frameworks
  • Experience with data drift detection tools and visualization libraries
  • Understanding of model performance metrics and drift detection methods
  • Familiarity with AI vulnerabilities and mitigation tools such as Adversarial Robustness Toolbox
  • Strong problem-solving, debugging, collaboration, communication, and attention-to-detail skills
  • Eligibility to obtain a Department of Homeland Security EOD clearance, including U.S. citizenship and a favorable background investigation
  • Familiarity with JavaScript or Go
  • Experience with LLM monitoring tools such as LangSmith or Helicone
  • Knowledge of GDPR or HIPAA compliance frameworks
  • Contributions to open-source MLOps projects or familiarity with MLOps/AIOps communities
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
1,176 Employees
Year Founded: 1993

What We Do

Entarian delivers science, engineering, and technology solutions for defense and federal civilian missions. The company integrates space systems, data modeling and analysis, cybersecurity, and secure enterprise and mission environments, enabling actionable insight from the enterprise to the tactical edge. Its work supports weather data, satellite operations, search and rescue, network infrastructure, and other mission-critical government programs for national security and public safety.

Similar Jobs

Capital One Logo Capital One

Lead Software Engineer

Fintech • Machine Learning • Payments • Software • Financial Services
Hybrid
McLean, VA, USA
55000 Employees
197K-225K Annually

Torc Robotics Logo Torc Robotics

Machine Learning Engineer

Artificial Intelligence • Automotive • Robotics • Software • Transportation
Remote or Hybrid
United States
500 Employees
132K-132K Annually

Torc Robotics Logo Torc Robotics

Software Engineer

Artificial Intelligence • Automotive • Robotics • Software • Transportation
Remote or Hybrid
United States
500 Employees
139K-167K Annually

TinyFish Logo TinyFish

MLOps Engineer

Artificial Intelligence
Remote or Hybrid
2 Locations

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Outpost Space Thumbnail
Aerospace • Defense
US
24 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account