As a Site Reliability Engineer (SRE), you will build and operate highly available, globally distributed advertising/monetization services. You will improve reliability, scalability, and operability through automation, observability, incident management, and sound engineering practices.
Key Responsibilities- Own reliability across the service lifecycle: design reviews, capacity planning, launch, deployment, operations, and continuous improvement.
- Build and operate highly available services across multiple regions/data centers; improve resilience, latency, and scalability.
- Develop automation and tooling to reduce toil (deployment, remediation, runbooks, self-healing) using scripting and software engineering best practices.
- Define and implement SLOs/SLIs/SLAs; create dashboards and alerting to track service health (availability, latency, errors, saturation).
- Lead sustainable incident response: triage, mitigation, root-cause analysis (RCA), and blameless postmortems with actionable follow-ups.
- Collaborate with software engineering, security, and compliance stakeholders to meet data governance and regulatory requirements.
- 3+ years of experience in SRE, DevOps, systems engineering, or production operations for large-scale services.
- Strong coding skills in one language: Python or Go or C++ (Java acceptable).
- Solid Linux/Unix fundamentals: processes, memory/CPU, filesystems, permissions, and troubleshooting.
- Networking fundamentals in cloud environments: TCP/IP, DNS, HTTP/HTTPS, load balancing, basic security concepts.
- SQL proficiency and experience with data workflows/ETL is a plus for ads/analytics-related systems.
- Strong communication, ownership mindset, and ability to work effectively across global teams.
- Experience supporting advertising, recommendation, or high-traffic consumer internet platforms.
- Hands-on experience with cloud platforms (AWS/GCP/Azure) and infrastructure-as-code (Terraform/Ansible).
- Experience with containers and orchestration (Docker, Kubernetes).
- Observability experience with tools such as Prometheus, Grafana, ELK/Splunk, OpenTelemetry.
- Experience operating large data systems (streaming, distributed storage/compute) and performance tuning.
Skills Required
- 3+ years experience in SRE, DevOps, systems engineering, or production operations for large-scale services
- Strong coding skills in one language: Python, Go, C++ (Java acceptable)
- Solid Linux/Unix fundamentals (processes, memory/CPU, filesystems, permissions, troubleshooting)
- Networking fundamentals in cloud environments (TCP/IP, DNS, HTTP/HTTPS, load balancing, basic security concepts)
- SQL proficiency and experience with data workflows/ETL
- Strong communication, ownership mindset, and ability to work across global teams
- Experience supporting advertising, recommendation, or high-traffic consumer internet platforms
- Hands-on experience with cloud platforms (AWS/GCP/Azure) and infrastructure-as-code (Terraform/Ansible)
- Experience with containers and orchestration (Docker, Kubernetes)
- Observability experience (Prometheus, Grafana, ELK/Splunk, OpenTelemetry)
- Experience operating large data systems (streaming, distributed storage/compute) and performance tuning
What We Do
Two95 International Inc., is a global technology firm specializing in enterprise solutions that evolves over BPM, Mobility, Cloud, Analytics, E-commerce & Social Business. Our client base includes several Fortune 500 and mid-market companies across industries and varying geographies. With vast knowledge and knowhow of 20 years in the IT field, we have been chosen as INC500 fastest growing company in North America in 2013. With the accolade of being ranked 11th in Human Resources by INC500, we have also been nominated as the 3rd fastest growing company in South Jersey by SJBM. We are ranked among the Top 20 IT Companies in New Jersey based on the year-on-year growth for the last 3 years. With a seasoned team of highly qualified personnel, our offices are located in New Jersey, Canada and India.








