Senior Software Engineer
About Us:
At Fullbay, our mission is simple — to create safer roads for our families and yours. As leaders in the heavy-duty repair industry, we power shops with technology that helps them run smarter and more efficiently. As an AI-First company, we invite artificial intelligence to eliminate friction, spark innovation, and drive efficiencies in every conversation— for our teams and our customers.
Position Overview:
The Senior Software Engineer is a key individual contributor focused on developer support operations and observability across Fullbay's cloud platform. This role exists to make internal engineers faster, more confident, and better equipped to build and operate production systems on Fullbay Next. Working closely with the Sherpa platform team and engineering teams across the organization, this engineer builds and maintains the telemetry infrastructure, tooling, and standards that give developers clear visibility into system behavior, cost, and reliability. Leveraging deep expertise in AWS-native architectures, distributed systems, and AI-assisted development practices, the Engineer designs observability solutions that are easy to adopt, hard to misuse, and actionable by default. The role drives best practices for instrumentation, infrastructure-as-code, and operational readiness across a 100% AWS microservices environment.
Primary Duties & Responsibilities:
- Design, build, and maintain observability infrastructure including distributed tracing, structured log aggregation, metrics pipelines, and alerting systems across Fullbay Next microservices using CloudWatch, OpenTelemetry, and AWS-native services
- Build and operate AI-assisted anomaly detection capabilities that help engineering teams identify and resolve production issues faster, reducing mean time to detection and resolution across the platform
- Develop internal developer tooling and self-service operational dashboards that surface real-time system health, cost visibility, and reliability signals, empowering engineers to own their services in production
- Establish and govern Terraform standards for provisioning observability infrastructure, writing reusable modules and reviewing infrastructure-as-code contributions across engineering teams
- Define and enforce instrumentation standards so that every service on Fullbay Next emits consistent, high-quality telemetry from day one, making log aggregation, tracing, and alerting reliable and low-maintenance across the platform
- Drive FinOps visibility by building cost and usage dashboards that give engineering and leadership clear, actionable insight into AWS spend at the service level
- Establish SLI, SLO, and error budget frameworks and build the systems that measure and report on them, enabling teams to make data-driven reliability tradeoffs
- Collaborate with Sherpa platform team members to integrate observability tooling into the internal developer platform, making provisioning, deployment, and operational monitoring seamless for every engineering team
- Investigate and evaluate emerging observability technologies and AI tooling, making informed build-vs-buy recommendations that align with Fullbay's AWS-first, serverless-preferred architecture
- Adhere to all confidentiality and compliance regulations
- Perform other duties as assigned
Minimum Education & Work Experience:
- This job requires at least 7-10 years of experience in Software Design and Development, with significant depth in observability engineering, platform engineering, or developer support operations in a cloud-native environment.
Key Skills and Qualifications:
- Hands-on experience with CloudWatch and OpenTelemetry for distributed tracing, structured logging, metrics collection, and alerting across microservices architectures
- Strong proficiency in Java, TypeScript/Node.js, and Python; experience building production services on AWS Lambda, ECS, and DynamoDB in a 100% AWS environment
- Terraform expertise: writing, maintaining, and governing infrastructure-as-code for observability systems and developer platform components
- Experience building AI-assisted operational tooling, including anomaly detection, log analysis, or intelligent alerting systems that reduce manual triage burden
- Deep familiarity with FinOps principles and AWS cost visibility tooling, with the ability to build service-level cost dashboards that drive accountability across engineering teams
- Awareness of SRE best practices including SLI/SLO definition, error budgets, and reliability-first service design within a microservices platform
- Strong ability to define and communicate instrumentation standards, developer platform conventions, and operational readiness requirements across engineering teams of varying experience levels
Physical Demands and Work Environment:
The physical demands described here are representative of those that must be met by an employee to perform the essential functions of this job successfully. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions.
- Regularly required to sit at a desk in front of a computer and use hands to finger, handle, or feel objects, tools, or controls (including a computer keyboard and operating a telephone), lift and/or move up to 10 pounds.
- Frequently requires the use of hands and arms for reaching, as well as the ability to walk and communicate effectively through speaking and listening.
- Specific vision abilities required by this position include close vision, color vision, and the ability to adjust focus.
- Noise level in the work environment is usually moderate.
- Type on a computer keyboard and look at a computer monitor, and operate a cell phone or a computer-based phone.
Skills Required
- 7-10 years of experience in software design and development with significant depth in observability, platform engineering, or developer support operations in a cloud-native environment.
- Hands-on experience with CloudWatch and OpenTelemetry for distributed tracing, structured logging, metrics collection, and alerting across microservices.
- Strong proficiency in Java, TypeScript/Node.js, and Python; experience building production services on AWS Lambda, ECS, and DynamoDB in a 100% AWS environment.
- Terraform expertise: writing, maintaining, and governing infrastructure-as-code for observability systems and developer platform components.
- Experience building AI-assisted operational tooling (anomaly detection, log analysis, intelligent alerting) to reduce manual triage.
- Deep familiarity with FinOps principles and AWS cost visibility tooling; ability to build service-level cost dashboards.
- Awareness of SRE best practices including SLI/SLO definition, error budgets, and reliability-first service design.
- Ability to define and communicate instrumentation standards, developer platform conventions, and operational readiness across engineering teams.
What We Do
Fullbay is cloud-based shop management software built specifically for heavy duty repair shops. We are changing the industry so shop owners and their technicians can get more done in less time and have a life outside the shop. For career opportunities, please visit fullbay.com/careers.

.png)






