Comet is building the development platform for teams who want to ship robust, reliable, and responsible AI applications. Opik, our open source LLM evaluation framework, has quickly become one of the most popular tools in the space. Our experiment management platform is used by data scientists at companies like Uber, Netflix, and Etsy. Tens of thousands of researchers, engineers, academics, and hobbyists use Comet every day to build the future of AI.
Working at Comet will give you access to the most exciting work being done in all areas of machine learning. Some of the top researchers and companies working on self-driving cars, drug discovery, particle research, diffusion models, and LLMs use Comet every day. Your work has the potential to accelerate the development of some of the most impactful technology in the world, and you will be doing it alongside a team of passionate, caring individuals. If that sounds exciting, Comet is the right place for you.
Comet is backed by more than $63 million in venture capital funding and powers some of the best machine-learning teams in the world, including Netflix, Uber, Etsy, and Mobileye. We are a remote-first company with offices in New York City (USA) and Tel Aviv (Israel).
Comet is looking for an experienced Deployment Engineer (located in the East or Central USA Only) to help our customers successfully deploy, operate, and troubleshoot Comet in their environments.
This is a highly customer-facing engineering role. You will work directly with customers to understand their infrastructure and deployment requirements, lead deployments, troubleshoot complex issues, and drive problems through to resolution.
You will also contribute to the engineering and automation behind our deployment solutions, including Kubernetes, Helm, Terraform, cloud infrastructure, tooling, and documentation.
What You’ll Do:- Work directly with customers to understand their infrastructure and deployment requirements.
- Deploy and maintain Comet across Kubernetes and cloud environments.
- Troubleshoot complex infrastructure and deployment issues across Kubernetes, networking, cloud services, databases, and application configuration.
- Own customer issues from investigation through resolution, collaborating with Engineering, Support, Customer Success, and other teams when needed.
- Develop and improve Helm charts, Terraform, deployment tooling, automation, and documentation.
- Participate in technical customer calls and explain complex topics clearly.
- Use AI tools for troubleshooting, development, automation, and investigation while validating results and maintaining engineering judgment.
- Proven customer-facing technical experience is essential.
- Proven experience using AI tools in real engineering workflows, with the ability to critically evaluate and validate their output.
- Strong hands-on experience with Kubernetes, Helm and containerized applications.
- Production experience with cloud infrastructure, preferably AWS.
- Experience with Infrastructure as Code, preferably Terraform.
- Strong Linux, networking, and infrastructure troubleshooting skills.
- Experience troubleshooting production systems using logs, metrics, and monitoring/observability tools.
- Excellent written and verbal communication skills.
- You take ownership and proactively move problems forward.
- You are passionate about troubleshooting unfamiliar systems and looking for root causes rather than quick workarounds.
- You can work independently while knowing when to collaborate or ask for help.
- You communicate progress, blockers, and decisions clearly - especially in a remote environment.
- You are comfortable with ambiguity and changing priorities.
- You work well across teams and care about both the customer experience and engineering quality.
- Enterprise or self-hosted software experience.
- Multi-cloud or on-premises deployment experience.
- CI/CD and GitOps experience.
- MLOps or AI infrastructure experience.
- Software development experience.
- Startup experience.
- Competitive base salary - $150K-200K based on proven experience, skills and location.
- Competitive benefits package.
- Flexible working hours and remote work options.
- Opportunities for professional growth and development.
- A collaborative and innovative work environment.
- The chance to work with cutting-edge technologies and projects.
- This role will be fully remote , located in the East or Central USA Only, working with a global team (large presence in the US, Europe and Tel Aviv) – some flexibility with work hours is required.
Please note: We do not accept unsolicited resumes or candidate submissions from recruitment agencies for this position.
Comet is an equal-opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees without regard to race, religion, color, sex, gender identity, gender expression, sexual orientation, national origin, ancestry, citizenship status, uniform service member status, marital status, pregnancy, age, medical condition, physical or mental disability, genetic information/characteristics, and any other characteristic protected by State or Federal law.
Skills Required
- Proven customer-facing technical experience
- Experience using AI tools in real engineering workflows and validating their output
- Strong hands-on experience with Kubernetes, Helm, and containerized applications
- Production experience with cloud infrastructure, preferably AWS
- Experience with Infrastructure as Code, preferably Terraform
- Strong Linux, networking, and infrastructure troubleshooting skills
- Experience troubleshooting production systems using logs, metrics, and monitoring or observability tools
- Excellent written and verbal communication skills
- Enterprise or self-hosted software experience
- Multi-cloud or on-premises deployment experience
- CI/CD and GitOps experience
- MLOps or AI infrastructure experience
- Software development experience
- Startup experience
What We Do
Comet is a meta machine learning platform designed to help AI practitioners and teams build reliable machine learning models for real-world applications by streamlining and connecting the machine learning model lifecycle. By leveraging Comet, users can employ machine learning experiment tracking to track, compare, explain and reproduce their models. Backed by thousands of users and multiple Fortune 100 companies, Comet provides insights and data to build better, more accurate AI models while improving productivity, collaboration and visibility across teams.








