Site Reliability Engineer

Posted 14 Hours Ago
Be an Early Applicant
Hiring Remotely in Sri Lanka
Remote
Junior
Food • Logistics
The Role
Designs and operates reliable, scalable software and platform services. Responsibilities include building automation and internal tools, improving observability, defining SLIs and SLOs, troubleshooting production systems, supporting incident response, enhancing deployment safety, and contributing to capacity planning, disaster recovery, resilience testing, and architecture reviews. The role partners with engineering teams and participates in on-call rotations to reduce toil and improve service reliability.
Summary Generated by Built In
JOB DESCRIPTION
Software Site Reliability Engineer

About Sysco LABS:

Sysco LABS is the Global In-House Center of Sysco Corporation (NYSE: SYY), the world’s largest foodservice company. Sysco ranks 55th in the Fortune 500 list and is the global leader in the trillion-dollar foodservice industry.


Sysco operates 333 distribution centers across 10 countries, with 75,000 colleagues serving approximately 670,000 customer locations, including restaurants, healthcare and educational facilities, lodging establishments, entertainment venues and more. For fiscal year 2026, which ended June 27, 2026, the company generated sales of more than $84 billion.


Sysco LABS Sri Lanka delivers the technology that powers Sysco’s end-to-end operations.


Sysco LABS’ enterprise technology is present in the end-to-end foodservice journey, enabling the sourcing of food products, merchandising, storage and warehouse operations, order placement and pricing algorithms, the delivery of food and supplies to Sysco’s global network and the in-restaurant dining experience of the end-customer.


The Opportunity


Join the Sysco Commercial Technology (CT) Site Reliability Engineering team as a Software Site Reliability Engineer, where you will help improve the reliability, scalability, performance, security, and operational excellence of Sysco's Commercial Technology ecosystem. Our team is responsible for delivering end-to-end reliability across the technology stack from cloud infrastructure and platform services to APIs, applications, and customer-facing digital experiences supporting enterprise API platforms, digital commerce, B2B integrations, and other business-critical systems across the CT landscape.


This is a Software SRE role focused on applying software engineering principles to solve reliability challenges at scale. You will design and build automation, develop internal tools and platforms, enhance observability, reduce operational toil, and deliver engineering solutions that strengthen the reliability and resilience of distributed systems.

Working closely with product, platform, and infrastructure engineering teams, you will contribute to building highly available, scalable, and resilient services while embedding reliability throughout the software development lifecycle.


Responsibilities:

  • Own the reliability, scalability, performance, and operational excellence of one or more services or platform components across the Commercial Technology ecosystem.
  • Apply software engineering principles to design and build solutions that improve reliability, resilience, and engineering productivity.
  • Design, develop, and maintain automation, internal tools, self-service capabilities, and reliability engineering workflows.
  • Define, implement, and improve SLIs, SLOs, observability standards, dashboards, alerts, and operational metrics.
  • Partner with product, platform, and infrastructure engineering teams to improve service architecture, production readiness, deployment safety, and operational excellence.
  • Troubleshoot complex production issues using application code, logs, metrics, traces, APIs, databases, and infrastructure telemetry to identify root causes and implement long-term solutions.
  • Participate in incident response, postmortems, and reliability reviews, ensuring follow-up actions are implemented through engineering improvements.
  • Improve deployment automation, release validation, rollback strategies, and production readiness.
  • Contribute to capacity planning, performance optimization, disaster recovery, and resilience testing.
  • Participate in architecture and design reviews to ensure systems are reliable, scalable, observable, secure, and operationally efficient.
  • Build and enhance shared engineering tools, libraries, and platform capabilities that improve developer productivity and service reliability.
  • Participate in an on-call rotation, using operational insights to continuously improve automation and reduce operational toil.

Requirements:

  • 2+ years of experience in Site Reliability Engineering, Software Engineering, Platform Engineering, Production Engineering, DevOps, or a related engineering role.
  • Strong software engineering skills, with experience designing, developing, testing, and operating production-grade software.
  • Proficiency in one or more programming languages, such as Java, Go, Python, or JavaScript/TypeScript.
  • Ability to read, understand, and debug application code to identify systemic issues and implement long-term engineering improvements.
  • Solid understanding of distributed systems, cloud-native architectures, microservices, APIs, databases, messaging systems, caching, networking, and system design.
  • Experience supporting large-scale distributed systems, enterprise SaaS solutions, digital commerce platforms, or other high-traffic environments.
  • Experience applying Site Reliability Engineering principles, including SLIs, SLOs, error budgets, incident management, blameless postmortems, automation, and toil reduction.
  • Experience troubleshooting production systems using application code, telemetry, and infrastructure diagnostics.
  • Hands-on experience with observability practices, including metrics, logs, traces, and profiling, using platforms such as Datadog, Prometheus, Grafana, OpenTelemetry, Splunk, or ELK.
  • Experience with AWS, Azure, or GCP and cloud-native technologies such as containers and Kubernetes.
  • Experience with CI/CD, Infrastructure as Code, and platform engineering tools such as Terraform, Helm, Argo CD, Jenkins, or GitHub Actions.
  • Experience developing internal developer platforms, automation frameworks, or reliability engineering tools that improve reliability and engineering productivity.
  • Experience with service meshes, API gateways, messaging systems, caching, database reliability, or resilience engineering.
  • Experience with AIOps, AI-assisted operations, or modern observability platforms.
  • Strong analytical, problem-solving, and systems-thinking skills.
  • Excellent communication and collaboration skills, with the ability to work effectively across cross-functional engineering teams.
  • Demonstrated ownership, curiosity, and a continuous-improvement mindset.

Benefits: 

  • Performance-based annual bonus  
  • Performance rewards and recognition  
  • Agile Benefits - special allowances for Health, Wellness & Academic purposes  
  • Paid birthday leave 
  • Team engagement allowance  
  • Comprehensive Health & Life Insurance Cover - extendable to parents and in-laws  
  • Hybrid work arrangement  

  Sysco LABS is an Equal Opportunity Employer.

Skills Required

  • 2+ years of experience in Site Reliability Engineering, Software Engineering, Platform Engineering, Production Engineering, DevOps, or a related engineering role.
  • Strong software engineering skills, including designing, developing, testing, and operating production-grade software.
  • Proficiency in one or more programming languages such as Java, Go, Python, or JavaScript/TypeScript.
  • Ability to read, understand, and debug application code.
  • Understanding of distributed systems, cloud-native architectures, microservices, APIs, databases, messaging systems, caching, networking, and system design.
  • Experience supporting large-scale distributed systems, enterprise SaaS solutions, digital commerce platforms, or high-traffic environments.
  • Experience applying SRE principles, including SLIs, SLOs, error budgets, incident management, blameless postmortems, automation, and toil reduction.
  • Experience troubleshooting production systems using application code, telemetry, and infrastructure diagnostics.
  • Hands-on observability experience with metrics, logs, traces, and profiling using tools such as Datadog, Prometheus, Grafana, OpenTelemetry, Splunk, or ELK.
  • Experience with AWS, Azure, or GCP and cloud-native technologies such as containers and Kubernetes.
  • Experience with CI/CD, Infrastructure as Code, and tools such as Terraform, Helm, Argo CD, Jenkins, or GitHub Actions.
  • Experience developing internal developer platforms, automation frameworks, or reliability engineering tools.
  • Experience with service meshes, API gateways, messaging systems, caching, database reliability, or resilience engineering.
  • Experience with AIOps, AI-assisted operations, or modern observability platforms.
  • Strong analytical, problem-solving, and systems-thinking skills.
  • Excellent communication and cross-functional collaboration skills.
  • Demonstrated ownership, curiosity, and a continuous-improvement mindset.

Sysco Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Sysco and has not been reviewed or approved by Sysco.

  • Healthcare Strength — Multiple national medical plan options with telehealth, behavioral health resources, and targeted programs indicate broad coverage and support. Preventive care access and ancillary offerings (dental, vision, Rx advocacy) further reinforce the package.
  • Retirement Support — A 401(k) with automatic company contributions plus a match, alongside an employee stock purchase plan, underscores solid retirement support. At union locations, enhanced pension terms add to perceived long‑term value.
  • Pay Growth & Progression — Recent collective bargaining outcomes with substantial wage increases demonstrate meaningful pay progression where contracts apply. In high‑volume markets, incentive structures can amplify earnings beyond base rates.

Sysco Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Houston, TX
24,120 Employees

What We Do

Sysco is the global leader in selling, marketing and distributing food and related products to customers who prepare meals away from home. This includes restaurants, healthcare and educational facilities, lodging establishments, entertainment venues, and more. Sysco operates almost 340 distribution centers, in over 10 countries, with 76,000 colleagues serving approximately 730,000 customer locations. The company generated sales of more than $81 billion in fiscal year 2025 that ended June 28, 2025. As the world’s largest food-away-from-home distributor, Sysco offers customized supply chain solutions, bespoke specialty product offerings, and culinary support to drive customers to innovate and optimize their operations. We act as a trusted business partner to our customers, helping them grow through our industry-leading portfolio that includes fresh produce, premium proteins, specialty products, sustainably focused items, equipment and supplies, and innovative culinary solutions. For more information, visit www.sysco.com. For important news and key information for Sysco investors, visit the Investor Relations section of the company’s website at investors.sysco.com.

Similar Jobs

Remote
Sri Lanka
24120 Employees
Remote or Hybrid
Sri Lanka
29811 Employees

GXA Logo GXA

Senior Systems Engineer

Information Technology
Remote
4 Locations
40 Employees

IFS Logo IFS

Senior Business Analyst

Information Technology • Software
Remote or Hybrid
Western Province, LKA
6788 Employees

Similar Companies Hiring

Tastewise Thumbnail
Artificial Intelligence • Big Data • Food • Retail • Software • Generative AI • Big Data Analytics
NYC, NYC
120 Employees
Axle Health Thumbnail
Artificial Intelligence • Healthtech • Information Technology • Logistics
Santa Monica, CA
25 Employees
Amalgamated Sugar Thumbnail
Food • Greentech • Agriculture • Industrial • Manufacturing
Boise, Idaho
768 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account