Senior SRE & Monitoring Developer

Sorry, this job was removed at 06:52 p.m. (UTC) on Tuesday, Sep 08, 2026
Be an Early Applicant
2 Locations
Remote or Hybrid
Senior level
Automotive
The Role
Own and optimize Elastic-based observability for a mission-critical 3DX PLM platform. Design monitoring, alerting, health checks, dashboards, and custom Vega visualizations; analyze performance and capacity; manage Elastic clusters, ILM, Fleet, APM, access controls, and query optimization. Automate operational tasks using Python, Bash, or Go, collaborate with infrastructure and engineering teams, and maintain documentation, runbooks, and incident knowledge.
Summary Generated by Built In

Our Product Development (PD) Platform Operations team is responsible for the 24x7 uptime, stability, and performance of our internally customized deployment of the 3DX PLM (Product Lifecycle Management) platform — including its underlying infrastructure, network, and application layers. This platform is mission-critical to our engineering and manufacturing organizations.

We are looking for a Senior SRE & Monitoring Developer who combines deep hands-on Elastic Stack expertise with strong site reliability engineering practices to help us proactively detect, diagnose, and prevent issues before they impact our internal customers.

Responsibilities

Elastic Platform Ownership: Maintain deep expertise in Elastic, including cluster management, performance tuning, index/shard optimization, and Fleet & APM management. Conduct deep dives into Elastic query performance and resource consumption, specific to the Dassault 3DX product line. Manage index lifecycle policies (ILM) to balance performance, cost, and retention requirements

Proactive Monitoring & Alerting: Design, implement, and manage comprehensive monitoring and alerting systems across our platforms, with a specific focus on Elastic clusters. Define key metrics (SLIs), thresholds (SLOs), and escalation procedures to proactively identify and address potential issues before they become incidents. Develop and maintain a suite of automated health checks for critical endpoints, APIs, infrastructure components, and network paths

Author advanced ES|QL (Elasticsearch Query Language) queries to aggregate, transform, and analyze log, metric, and trace data for performance analysis, capacity planning, and root cause investigations.

Utilize KQL (Kibana Query Language) to create efficient  search filters, Discover saved views, dashboard controls, and alerting rule conditions. Translate ad-hoc operational and diagnostic questions into performant ES|QL pipelines that can be operationalized into Kibana dashboards or scheduled alerts.

Design and build advanced custom Kibana visualizations using Vega and Vega-Lite for complex use cases beyond out-of-the-box Kibana Lens capabilities (e.g., SLO burn-rate tracking, multi-layered service performance visuals, topology maps). Develop and maintain operations, executive, and incident dashboards combining standard Kibana panels with custom Vega visual specifications.

Performance Analysis & Optimization: Conduct regular performance analysis of platforms, identifying bottlenecks and implementing optimizations to improve responsiveness, scalability, and resource utilization. This includes deep dives into Elastic query performance & resource consumption – specific to the Dassault 3DX product line

Security & Compliance: Apply appropriate access control (e.g., RBAC in Elastic/Kibana) and follow data handling/security policies relevant to a manufacturing/engineering data environment

Automation & Tooling: Build and maintain scripts/tools (Python, Bash, or Go) to streamline health checks, monitoring integrations, and routine operational tasks

Collaboration & Enablement: Work closely with infra teams to ensure the reliability and scalability of applications that interact with Elastic. Provide guidance to engineering teams on Elastic best practices, query optimization, and observability instrumentation

Documentation & Knowledge Sharing: Create and maintain comprehensive documentation for monitoring systems, alerting rules, escalation procedures, and troubleshooting guides. Contribute to team runbooks and post-incident reviews to build institutional knowledgeQualificationsMust-Have
  • 7 to 10 years of experience as a Site Reliability Engineer, DevOps Engineer, or similar role with a strong focus on monitoring, observability, and incident management

  • Experience designing, implementing, and managing monitoring and alerting systems using industry-standard tools (ELK/Elastic Stack, Dynatrace)

  • Experience working with a PLM platform such as Dassault Systèmes 3DX/ Siemens Teamcenter

  • Strong scripting or programming skills (Python, Bash, or Go) for automating tasks, developing health checks, and integrating monitoring tools

  • Excellent problem-solving skills, a proactive mindset, and the ability to remain calm and effective under pressure during incidents

  • Strong communication and collaboration skills, with the ability to work effectively across technical and non-technical teams

Highly Desirable
  • Experience with infrastructure-as-code or Pipeline as a code tools (Ansible, Terraform)

  • Experience with cloud platforms (GCP, Azure) and cloud-native monitoring/logging services

  • Experience working within an Agile/DevOps environment

  • Elastic Certified Engineer certification or equivalent

  • Experience in manufacturing or engineering/CAD-adjacent industries

Skills Required

  • 7 to 10 years of experience as a Site Reliability Engineer, DevOps Engineer, or similar role focused on monitoring, observability, and incident management
  • Experience designing, implementing, and managing monitoring and alerting systems using ELK/Elastic Stack and Dynatrace
  • Experience working with a PLM platform such as Dassault Systemes 3DX or Siemens Teamcenter
  • Strong scripting or programming skills in Python, Bash, or Go
  • Excellent problem-solving skills and ability to remain calm and effective during incidents
  • Strong communication and collaboration skills across technical and non-technical teams
  • Experience with infrastructure-as-code or pipeline-as-code tools such as Ansible or Terraform
  • Experience with cloud platforms such as GCP or Azure and cloud-native monitoring/logging services
  • Experience working within an Agile/DevOps environment
  • Elastic Certified Engineer certification or equivalent
  • Experience in manufacturing or engineering/CAD-adjacent industries

Ford Motor Company Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Ford Motor Company and has not been reviewed or approved by Ford Motor Company.

  • Healthcare Strength Medical, dental, and vision coverage start on day one with options that include zero-premium plans, free mental health support, and wellness resources. For represented hourly employees, health plans are described as low-cost with strong coverage value.
  • Retirement Support A 401(k) with company match and additional company contributions is available from day one, alongside life and disability coverage. Pension eligibility in certain situations and financial-planning support reinforce long‑term security.
  • Parental & Family Support Paid parental leave, fertility, surrogacy, and adoption benefits, plus a ramp‑up program for returning parents, reflect a family‑focused package. Flexible Family Care days and generous time‑off options help address short‑term caregiving and personal needs.

Ford Motor Company Insights

Similar Jobs

Atlassian Logo Atlassian

Infrastructure Engineer

Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Remote
India
11000 Employees

Zapier Logo Zapier

Director, India

Artificial Intelligence • Productivity • Software • Automation
Remote
India
800 Employees

Micron Technology Logo Micron Technology

Shift Technician - Gas Operations (Facilities)

Artificial Intelligence • Hardware • Information Technology • Machine Learning
Remote
Gujarat, IND
45000 Employees

Micron Technology Logo Micron Technology

Module Planning Manager

Artificial Intelligence • Hardware • Information Technology • Machine Learning
Remote
Gujarat, IND
45000 Employees
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Dearborn, MI
175,633 Employees
Year Founded: 1903

What We Do

Ford is a global company with shared ideals and a deep sense of family. From our earliest days as a pioneer of modern transportation, we have sought to make the world a better place – one that benefits lives, communities and the planet. We are here to provide the means for every person to move and pursue their dreams, serving as a bridge between personal freedom and the future of mobility. In that pursuit, our 186,000 employees around the world help to set the pace of innovation every day.

Similar Companies Hiring

Cox Enterprises Thumbnail
Artificial Intelligence • Automotive • Greentech • Information Technology • Machine Learning • Software • Cybersecurity
Atlanta, Georgia
30000 Employees
UL Solutions Thumbnail
Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Chicago, IL
15000 Employees
HERE Technologies Thumbnail
Artificial Intelligence • Automotive • Computer Vision • Information Technology • Internet of Things • Logistics • Software
Amsterdam, NL
6000 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account