**ACTIVE TS/SCI SECURITY CLEARANCE REQUIRED**
We're looking for a Senior Application Monitoring Engineer to own and evolve our application observability strategy. You'll be the go-to expert for keeping critical business applications visible, healthy, and performant — designing monitoring architecture, driving down mean time to detection and resolution, and mentoring other engineers on observability best practices. This is a largely autonomous role for someone who thinks proactively about failure modes and treats monitoring as a product, not an afterthought.
What You'll Do- Design, build, and maintain end-to-end application performance monitoring using platforms like Dynatrace, instrumenting applications and services so that the right signals surface before customers notice a problem
- Define and maintain SLIs, SLOs, and error budgets in partnership with engineering and product teams, and build dashboards and alerting that reduce noise while catching what actually matters
- When incidents happen, you'll lead or support triage and root-cause analysis, using distributed tracing and APM data to pinpoint issues across complex, distributed systems.
- Drive continuous improvement of the observability stack itself — evaluating new tools, refining alert thresholds, and reducing alert fatigue across the engineering organization
- Serve as a technical mentor, helping other engineers build monitoring into their own services from the start rather than bolting it on later
- Partner closely with DevOps/SRE, infrastructure, and application development teams to make sure monitoring coverage keeps pace with new deployments and architectural changes
- 5+ years of experience in application monitoring, observability, or site reliability engineering, with hands-on expertise in at least one major APM platform (Datadog, New Relic, or Dynatrace) — deep familiarity with more than one is a strong plus
- Experience configuring alerts and adjusting thresholds
- Solid grasp of distributed systems, microservices architecture, and how to trace a request across services, containers, and cloud infrastructure
- Experience with scripting or programming (Python, Bash, or similar) to automate monitoring configuration and build custom integrations
- Experience with cloud platforms (AWS, Azure, or GCP) and containerized environments (Docker, Kubernetes) is expected
- Understanding of the fundamentals of APIs, databases, and networking well enough to diagnose issues across the full stack
- You've participated in on-call rotations and incident response processes, and ideally have exposure to complementary tools like Splunk, ELK, Grafana, or Prometheus
- Communicate clearly under pressure, can explain a complex incident to both engineers and non-technical stakeholders, and take genuine ownership of the systems you monitor
- Familiarity with CI/CD pipelines and infrastructure-as-code (Terraform, CloudFormation, etc.) is a plus
- Relevant certifications (Datadog, New Relic, AWS/Azure/GCP)
Additional Information
- All your information will be kept confidential according to EEO guidelines.
- Compensation is unique to each candidate and relative to the skills and experience they bring to the position. The salary range for this position is typically $150-175k. This does not guarantee a specific salary as compensation is based upon multiple factors such as education, experience, certifications, and other requirements, and may fall outside of the above-stated range.
- Highlights of our benefits include Health/Dental/Vision, 401(k) match, Accrued PTO, STD/LTD/Life Insurance, Referral Bonuses, professional development reimbursement, and more!
D2 Technical Services is committed to a merit-based recruitment process and encourages applications from all qualified individuals. As a Veteran-Owned Small Business, we particularly welcome applications from veterans who have the requisite skills and experience. Job applicants that are interested in one of our openings and may require a reasonable accommodation to participate in the job application or interview process, should contact us to request an accommodation.
Skills Required
- Active Top Secret/Sensitive Compartmented Information (TS/SCI) security clearance
- 5+ years of experience in application monitoring, observability, or site reliability engineering
- Hands-on expertise with a major APM platform such as Datadog, New Relic, or Dynatrace
- Experience configuring alerts and adjusting thresholds
- Understanding of distributed systems, microservices architecture, and distributed tracing
- Scripting or programming experience with Python, Bash, or similar
- Experience with AWS, Azure, or GCP
- Experience with Docker and Kubernetes
- Understanding of APIs, databases, and networking for full-stack diagnosis
- Participation in on-call rotations and incident response processes
- Familiarity with CI/CD pipelines and infrastructure-as-code such as Terraform or CloudFormation
- Relevant Datadog, New Relic, AWS, Azure, or GCP certifications
What We Do
D2 Consulting provides services to the Federal Government focused in the following three area's We leverage D2 Consulting engineering, operations and governance best practices to efficiently and effectively deploy, maintain and continuously improve IT services and solutions. This includes not only engineering solutions such as VDI and providing direct support to operations but also deploying the tools to instrument and maintain Enterprise performance and availability (including ITSM and Enterprise Management Tools). Protect & Secure: This is our cyber security practice, which includes functions like audit and information assurance, as well as the traditional mechanics of securing IT systems and services. We have experience working with the Government to better manage risk to help move the accreditation process along. This is especially true with respect to adapting accreditation controls to a cloud environment. Cloud Migration and Data Center modernization/consolidation: This area is focused on cloud migration and adoption as well as data center consolidation. We work to ensure that infrastructure requirements are implemented correctly, according to best practices and in a timely manner.








