Senior Software Engineer, Machine Learning Infrastructure (Tinder LLC, West Hollywood, California)
Design, build, and maintain scalable machine learning (ML) infrastructure to support experimentation, training, deployment, and monitoring of ML models processing large-scale datasets with hundreds of billions of data points.
Develop and maintain robust, scalable infrastructure platforms that support the needs of machine learning engineers across multiple business units. Design, build, and maintain data processing and moderation pipelines that handle large data volumes and integrate with trust and safety workflows. Deploy and manage production ML systems using internal deployment tools and optimize compute and storage resources to ensure reliability, scalability, and cost efficiency. Design, develop, and maintain application programming interfaces (APIs), including REST, gRPC, and GraphQL, to support internal ML platform services and system integrations. Oversee deployment, monitoring, and performance of ML systems using observability tools to ensure compliance with technical specifications and service-level objectives. Develop and implement model evaluation, validation, and quality assurance processes, including A/B testing frameworks and automated evaluation systems, to ensure model accuracy, reliability, and performance. Design, develop, and maintain scalable ML platform systems and data infrastructure using distributed data technologies, including Apache Spark, Kafka, Flink, and Databricks, to support global data processing and analytics needs. Analyze ML infrastructure requirements across business units and design technical solutions within defined scalability, performance, and cost constraints. Support technical design and implementation of ML lifecycle infrastructure, including model training, serving, monitoring, feature stores, and evaluation systems, with an emphasis on platform engineering and self-service capabilities. Mentor and provide technical guidance to junior engineers on ML systems, backend systems, scalable data pipelines, production reliability, and deployment best practices. Participate in hiring activities by conducting technical interviews and providing input on candidate evaluations. Develop and maintain technical documentation, including system designs, operational guides, and internal knowledge bases. Design and optimize recommendation systems and moderation data pipelines, applying best practices for data versioning, feature management, and model evaluation. Implement and optimization of backend and ML services to ensure reproducibility, reliability, and operational stability. Design and optimize large-scale data pipelines and database systems to support efficient data access patterns for ML workflows. Collaborate with cross-functional teams, including software engineers, data engineers, and ML engineers, to support the development and deployment of ML-enabled product features. Design and maintain infrastructure supporting large language model (LLM) workloads. Analyze and resolve complex distributed systems issues affecting performance, scalability, reliability, and availability of high-traffic ML applications. Research and evaluate emerging ML infrastructure technologies and conduct proof-of-concept implementations to support architectural and technology decisions. Stay current with advances in ML infrastructure, distributed systems, and data engineering, and apply industry best practices to ongoing platform development. Telecommuting may be permitted. When not telecommuting must report to 8800 Sunset Blvd. West Hollywood, CA 90069. Up to 10% domestic travel for team meetings and on-site trainings. Salary: $190K - $246K per year.
MINIMUM REQUIREMENTS: Bachelor’s degree or its U.S. equivalent in Computer Science, Computer Engineering, or a related field, plus 5 years of professional experience as a Machine Learning Engineer, Site Reliability Engineer, or any occupation/position/job title performing ML infrastructure or backend software engineering.
In lieu of a Bachelor’s degree plus 5 years of experience, the employer will accept a Master’s degree or U.S. equivalent in Computer Science, Computer Engineering ,or related field, plus 3 years of professional experience as a Machine Learning Engineer, Site Reliability Engineer, or any occupation/position/job title performing ML infrastructure or backend software engineering.
Must also have experience in the following: 3 years of professional experience designing and implementing large-scale distributed ML platform systems, using big data technologies including Apache Spark, Apache Kafka, Apache Flink, or Databricks. 3 years of professional experience using multiple modern programming languages, including Python, Scala, Java, or Go, to develop ML platform systems, backend services, data
processing jobs, and automation tools supporting the ML lifecycle. 2 years of professional experience working with modern cloud platforms (including AWS, Azure, or GCP) and utilizing infrastructure-as-code practices, containerization tools (Docker on managed orchestration platforms including Amazon EKS or Amazon ECS), and monitoring systems based on Prometheus metrics and Grafana dashboards, including experience operating services backed by a timeseries metrics store including Grafana Mimir. 2 years of professional experience designing and building infrastructure for recommendation systems, moderation pipelines, or large language model (LLM) serving and deployment systems, including experience with modern ML serving frameworks including Ray Serve or Triton, and with LLM-serving. 2 years of professional experience in large-scale database design and optimization, and data pipeline performance tuning to support efficient data access patterns for ML workflows, including working with analytical storage systems including Delta Lake or data warehouses, including Redis, ValKey or DynamoDB. 1 year of professional experience leading technical initiatives across multiple engineering teams, including establishing platform ownership models, providing hands-on technical guidance, and driving adoption of shared ML infrastructure components including standardized GitOps pipelines, and modern model-serving platforms. 1 years of professional experience designing and implementing CI/CD automation pipelines and GitOps practices for ML infrastructure, using tools including Terraform, Terragrunt, Helm, and internal GitOps systems (including Scaffold) together with continuous integration systems (including Jenkins or Buildkite) to manage deployment strategies including canary releases, bluegreen deployments, and zerodowntime migrations of backend services.
CONTACT: Please email resume to: [email protected]. Must specify Ad Code SLLL in subject line.
Skills Required
- Bachelor's degree in Computer Science, Computer Engineering, or related field plus 5 years of professional experience as an ML Engineer, Site Reliability Engineer, or ML infrastructure/backend software engineer.
- Master's degree in Computer Science, Computer Engineering, or related field plus 3 years of professional experience as an ML Engineer, Site Reliability Engineer, or ML infrastructure/backend software engineer (alternative to Bachelor's + 5 years).
- 3 years professional experience designing and implementing large-scale distributed ML platform systems using Apache Spark, Apache Kafka, Apache Flink, or Databricks.
- 3 years professional experience using multiple modern programming languages (Python, Scala, Java, or Go) to develop ML platform systems, backend services, data processing jobs, and automation tools.
- 2 years professional experience with modern cloud platforms (AWS, Azure, or GCP), infrastructure-as-code practices, containerization (Docker) on managed orchestration (Amazon EKS or Amazon ECS), and monitoring systems (Prometheus, Grafana, Grafana Mimir).
- 2 years professional experience designing and building infrastructure for recommendation systems, moderation pipelines, or LLM serving and deployment systems, including ML serving frameworks such as Ray Serve or Triton.
- 2 years professional experience in large-scale database design and optimization and data pipeline performance tuning, including analytical storage systems (Delta Lake, data warehouses) and operational stores (Redis, ValKey, DynamoDB).
- 1 year professional experience leading technical initiatives across multiple engineering teams, establishing platform ownership models, and driving adoption of shared ML infrastructure and standardized GitOps pipelines.
- 1 year professional experience designing and implementing CI/CD automation pipelines and GitOps practices for ML infrastructure using Terraform, Terragrunt, Helm, internal GitOps systems (Scaffold), and CI systems (Jenkins or Buildkite) supporting canary, blue/green, and zero-downtime deployments.
- Experience designing, developing, and maintaining APIs (REST, gRPC, GraphQL) and implementing observability, monitoring, model evaluation, validation, A/B testing, and quality assurance processes for ML systems.
Match Group Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Match Group and has not been reviewed or approved by Match Group.
-
Retirement Support — Employer matching on retirement savings and access to an employee stock purchase program are highlighted as standout elements. Feedback suggests this creates strong long-term financial support as part of total rewards.
-
Parental & Family Support — Fully paid parental leave, fertility and family-forming support, and childcare resources are emphasized across materials. Feedback suggests these benefits meaningfully support different family needs.
-
Leave & Time Off Breadth — Generous PTO, numerous paid holidays, wellness and volunteer time, and other special leave options are consistently described. Feedback suggests the breadth of time-off options helps sustain work-life balance.
Match Group Insights
What We Do
Match Group is home to some of the world’s most popular dating and social discovery apps, including Tinder, Hinge, Match, and more. Match Group’s mission is to spark meaningful connections for every single person worldwide. Our diverse portfolio of apps enables connections across a diverse range of ages, genders, backgrounds, and dating goals. Our services are available in over 40 languages to users all over the world.
Why Work With Us
Match Group is united by a mission to help people find meaningful connections and fight loneliness. We’re not just building products. We’re creating friendships, marriages, and families around the world. Our culture is fueled by purpose, creativity, and a passion for what we can build together.
Gallery







