Key Responsibilities
- Design and manage scalable ML infrastructure on AWS using EKS, EC2, RDS, S3, and IAM-based access control.
- Build and maintain Kubernetes deployments for LLM and TTS inference using Helm, ArgoCD, and Prometheus/Grafana monitoring.
- Implement and optimize model serving pipelines using vLLM, SGLang, TensorRT, or similar frameworks for high-throughput inference.
- Develop CI/CD and MLOps automation for data versioning, model validation, and deployment (GitHub Actions, Jenkins, or AWS CodePipeline).
- Integrate OpenWebUI, Gradio, or similar UIs for user-facing model demos and internal evaluation tools.
- Collaborate with ML researchers to productize models — including TTS (e.g., ElevenLabs API), ASR (Whisper), and LLM-based chat systems.
- Ensure observability, cost optimization, and reliability of cloud resources across multiple environments.
- Contribute to internal tools for dataset curation, model monitoring, and retraining pipelines.
- Maintain infrastructure-as-code using Terraform and Helm charts for reproducibility and governance.
- Support real-time multimodal workloads (voice, text, vision) across inference clusters.
Academic Qualifications
- 4+ years of experience in MLOps, DevOps, or Cloud Infrastructure Engineering for ML systems.
- Strong proficiency in Kubernetes, Helm, and container orchestration.
- Experience deploying ML models via vLLM, SGLang, TensorRT, or Ray Serve.
- Proficiency with AWS services (EKS, EC2, S3, RDS, CloudWatch, IAM).
- Solid experience with Python, Docker, Git, and CI/CD pipelines.
- Strong understanding of model lifecycle management, data pipelines, and observability tools (Grafana, Prometheus, Loki).
- Excellent collaboration skills with ML researchers and software engineers.
Professional Experience – Preferred
- Extensive Experience with vLLM, K8s, Elevenlabs, Whisper, Gradio/OpenWebUI, or custom TTS/ASR model hosting.
- Familiarity with multi-GPU scheduling, NCCL optimization, and HPC cluster integration.
- Knowledge of security, cost management, and network policy in multi-tenant Kubernetes clusters and cloudflare systems.
- Prior work in LLM deployment, fine-tuning pipelines, or foundation model research.
- Exposure to data governance and responsible AI operations in research or enterprise settings.
Skills Required
- 4+ years experience in MLOps, DevOps, or Cloud Infrastructure Engineering for ML systems
- Proficiency with Kubernetes (EKS), Helm, and container orchestration
- Experience deploying ML models via vLLM, SGLang, TensorRT, or Ray Serve
- Proficiency with AWS services (EKS, EC2, S3, RDS, CloudWatch, IAM)
- Strong experience with Python, Docker, Git, and CI/CD pipelines
- Strong understanding of model lifecycle management, data pipelines, and observability tools (Grafana, Prometheus, Loki)
- Maintain infrastructure-as-code using Terraform and Helm charts
- Implement CI/CD and MLOps automation (GitHub Actions, Jenkins, or AWS CodePipeline)
- Excellent collaboration skills with ML researchers and software engineers
- Experience with ElevenLabs, Whisper, RVC, Gradio/OpenWebUI or custom TTS/ASR model hosting
- Familiarity with multi-GPU scheduling, NCCL optimization, and HPC cluster integration
- Knowledge of security, cost management, and network policy in multi-tenant Kubernetes clusters and Cloudflare
- Prior work in LLM deployment, fine-tuning pipelines, or foundation model research
- Exposure to data governance and responsible AI operations
What We Do
First a passion, then an idea transformed into success – when it comes to pioneering automation and digitalisation technology, the ifm group is the ideal partner. Since its foundation in 1969, ifm has developed, produced and sold sensors, controllers, software and systems for industrial automation and for SAP-based solutions for supply chain management and shop floor integration worldwide. As one of the pioneers of Industry 4.0, ifm develops and implements consistent solutions to digitalise the entire value chain “from sensor to ERP”. Today, the second-generation family-run ifm group has more than 8,750 employees and is one of the worldwide market leaders. The group combines the internationality and innovative strength of a growing group of companies with the flexibility and close customer contact of a medium-sized company.







