- 5+ Years experience, Implement, maintain, and optimize AI Platform capabilities deployed on Google Kubernetes Engine (GKE) in Google Cloud Platform (GCP), enabling product teams to build and deploy Generative AI applications and self-hosted models.
- Apply software and platform engineering best practices—including code quality, automated testing, CI/CD, GitOps, OTel telemetry, and documentation—to maintain reliable platform components and internal developer tooling.
- Collaborate directly with application, data science, and platform teams to collect feedback, troubleshoot integration issues, and implement cloud-native AI tools and observability standards.
- Stay current with advancements in AI infrastructure, GPU resource management, LLM serving engines, and cloud-native networking to continuously improve platform efficiency and developer experience.
- Build and integrate platform components, custom K8s resources, and telemetry pipelines using solid software design patterns within the broader GCP/GKE ecosystem.
- Contribute to team technical documentation, create reusable platform templates ("paved paths"), and assist other engineers in adopting platform standards.
- Execute on team technical roadmaps to improve platform reliability, lower operational overhead, and speed up delivery cycles for AI product features.
1. Skills and Experience Requirements1. AI Infrastructure & Generative AI Experience
- Self-Hosted LLM Serving: Hands-on experience deploying, configuring, and scaling self-hosted Large Language Models (LLMs) on GKE using inference engines such as vLLM, NVIDIA NIM, or SGLang. Basic understanding of Kubernetes GPU allocation, multi-GPU nodes, and model execution requirements.
- Cloud-Native AI Tools: Experience working with or integrating emerging cloud-native AI infrastructure tools and control planes such as Agentgateway or Kagent to support agentic workflows and request routing.
- AI Agent & LLM Observability (OpenTelemetry): Practical experience configuring OpenTelemetry (OTel) instrumentation and collectors for AI applications. Familiarity with OTel GenAI Semantic Conventions (gen_ai.*) to capture traces, metrics, token usage, and latency across agentic execution flows, tool calls, and model calls.
- GenAI Frameworks: Hands-on experience integrating with GenAI frameworks (e.g., LangGraph, LangChain, Google Agent Development Kit/ADK) and cloud services (e.g., Google Vertex AI, Google Agentspace, Gemini APIs).
- Production Deployment: Experience deploying and supporting production workload pipelines in Kubernetes with attention to latency, availability, and resource utilization.
2. Google Cloud & Cloud-Native Networking
- GKE & GCP Knowledge: Solid, practical experience deploying and managing workloads on Google Kubernetes Engine (GKE), Google Compute Engine (GCE), and related GCP infrastructure.
- Service Mesh: Practical experience operating and troubleshooting Istio Ambient Mode (or sidecar-based Istio transitioning to Ambient) for mTLS, traffic routing, and service-to-service communication.
- Cloud-Native Tooling: Strong familiarity with container runtime environments, GPU device plugins/operators on Kubernetes, OpenTelemetry Collectors (OTLP), GitOps workflows (e.g., ArgoCD, Flux), Infrastructure as Code (Terraform), and standard security practices.
3. Engineering & Domain Standards
- Software Engineering Practices: Practical understanding of design patterns, unit/integration testing, clean code principles, and writing maintainable code.
- Programming Skills: Strong proficiency in Python or Go for writing automation, custom tooling, scripts, or operators.
- Teamwork & Communication: Strong collaboration skills with the ability to write clear documentation, work across team boundaries, and explain technical setup to fellow engineers.
CME Group: Where Futures are Made
CME Group is the world’s leading derivatives marketplace. But who we are goes deeper than that. Here, you can impact markets worldwide. Transform industries. And build a career by shaping tomorrow. We invest in your success and you own it – all while working alongside a team of leading experts who inspire you in ways big and small. Problem solvers, difference makers, trailblazers. Those are our people. And we’re looking for more.
At CME Group, we embrace our employees' unique experiences and skills to ensure that everyone’s perspectives are acknowledged and valued. As an equal-opportunity employer, we consider all potential employees without regard to any protected characteristic.
Important Notice: Recruitment fraud is on the rise, with scammers using misleading promises of job offers and interviews to solicit money and personal information from job seekers. CME Group adheres to established procedures designed to maintain trust, confidence and security throughout our recruitment process. Learn more here.
Skills Required
- 5+ years of professional experience
- Hands-on experience deploying, configuring, and scaling self-hosted LLMs on GKE using vLLM, NVIDIA NIM, or SGLang
- Understanding of Kubernetes GPU allocation, multi-GPU nodes, and model execution requirements
- Experience with cloud-native AI infrastructure tools or control planes such as Agentgateway or Kagent
- Practical experience configuring OpenTelemetry instrumentation and collectors for AI applications
- Familiarity with OTel GenAI Semantic Conventions for traces, metrics, token usage, and latency
- Hands-on experience integrating GenAI frameworks such as LangGraph, LangChain, or Google ADK
- Experience with Google Vertex AI, Google Agentspace, or Gemini APIs
- Experience deploying and supporting production workloads in Kubernetes
- Practical experience deploying and managing workloads on GKE, GCE, and related GCP infrastructure
- Practical experience operating and troubleshooting Istio Ambient Mode or sidecar-based Istio
- Familiarity with container runtimes, Kubernetes GPU device plugins or operators, OpenTelemetry Collectors, GitOps, Terraform, and security practices
- Understanding of design patterns, unit and integration testing, clean code, and maintainable software
- Strong proficiency in Python or Go
- Strong collaboration, documentation, and technical communication skills
CME Group Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about CME Group and has not been reviewed or approved by CME Group.
-
Retirement Support — U.S. offerings include both a 401(k) and a company-funded cash-balance pension, strengthening long-term financial security. This dual-track structure is highlighted as a notable differentiator among private employers.
-
Leave & Time Off Breadth — PTO and holiday schedules are described as generous, with ample time off and carryover commonly highlighted. This breadth of leave meaningfully enhances perceived total rewards.
-
Flexible Benefits — A flexible, hybrid work model applies to many roles, increasing day-to-day usability of the package. Flexibility is framed as a standard feature rather than an exception.
CME Group Insights
What We Do
As the world's leading derivatives marketplace, CME Group (www.cmegroup.com) is where the world comes to manage risk. CME Group exchanges offer the widest range of global benchmark products across all major asset classes, including futures and options based on interest rates, equity indexes, foreign exchange, energy, agricultural commodities, metals, weather and real estate. CME Group brings buyers and sellers together through its CME Globex® electronic trading platform and its trading facilities in New York and Chicago. CME Group also operates CME Clearing, one of the world’s leading central counterparty clearing provider in the world, which offers clearing and settlement services for exchange-traded contracts, as well as for over-the-counter derivatives transactions through CME ClearPort®. These products and services ensure that businesses everywhere can substantially mitigate counterparty credit risk in both listed and over-the-counter derivatives markets.








