Senior Advisory Software Engineer

Posted 19 Days Ago
Be an Early Applicant
Pune, Mahārāshtra, IND
In-Office
Senior level
Information Technology • Logistics • Financial Services
The Role
Architect and operate AI-augmented platform engineering systems for autonomous monitoring, anomaly detection, incident response, remediation, CI/CD reliability, infrastructure lifecycle management, observability, and FinOps. Define SLOs, reliability metrics, disaster recovery, runbook automation, and human-escalation safeguards. Lead postmortems, embed reliability practices across the SDLC, and provide senior technical leadership while building agentic workflows with LLM APIs and frameworks such as LangGraph, AutoGen, and CrewAI.
Summary Generated by Built In

We’re hiring at Pitney Bowes, where top talent builds meaningful careers and lasting impact. We Move fast, Deliver excellence, and Win together…that’s The Pitney Bowes way. Here, how we work matters just as much as what we achieve.

We’re looking for people who:

  • Act with urgency, accountability, and purpose

  • Deliver high quality work with consistency and pride

  • Collaborate effectively and elevate those around them

  • Focus on outcomes that drive impact and growth

Job Description:

Join Pitney Bowes as Senior Advisory Software Engineer

Years of experience: 9-12 years

Job Location – Pune

Impact

As a Senior Advisory Software Engineer, you will operate at the intersection of platform engineering and AI-driven automation. This is not a traditional SRE role. You will design, build, and supervise agentic systems that detect anomalies, diagnose failures, execute remediation runbooks, and escalate intelligently — with minimal human intervention. You will architect the feedback loops that make our platform progressively self-healing. You will collaborate across engineering, product, and architecture to ensure our observability and incident response capabilities stay ahead of system complexity. Being a Senior Advisory SRE here means you think in systems, build in agents, and measure success in mean-time-to-no-action.

The Job

  • Architect and operate agentic systems for autonomous monitoring, anomaly detection, and self-healing — reducing mean-time-to-remediation without human-in-the-loop dependency for routine failure patterns.

  • Build software and agentic pipelines that manage platform infrastructure autonomously — from drift detection and remediation to capacity adjustment and incident triage.

  • Drive reliability engineering outcomes — SLO attainment, error budget governance, and deployment velocity — by embedding intelligence into the platform rather than adding human process overhead.

  • Measure and continuously optimize system performance using agent-driven telemetry analysis — identifying degradation patterns before they manifest as customer-impacting incidents.

  • Own CI/CD reliability across the SDLC — integrating agentic checks, automated rollback triggers, and intelligent deployment gates that act on signal, not on schedule.

  • Scope includes:

    • Agentic observability — context-aware monitoring with LLM-assisted signal interpretation

    • Intelligent alert design — dynamic thresholds, noise suppression, and automated triage routing

    • Runbook automation and agentic remediation — codifying institutional knowledge into executable, supervised agent workflows

    • Autonomous incident response — agent-led detection, diagnosis, and escalation with human override at defined severity thresholds

    • Infrastructure lifecycle management — provisioning, drift remediation, and cost optimization driven by policy-as-code and agent execution

    • End-to-end configuration, deployment, and patching — governed by automated validation pipelines, not manual checklists

    • Creating and maintaining GIT repo and pipelines

  • Communicate risks, system health, and automation outcomes clearly to engineering leadership and cross-functional stakeholders — translating agent behaviour and reliability signals into business-readable insight.

  • Define and continuously refine the observability strategy — what to monitor, how to act on it, and how to suppress noise programmatically. Drive adoption of agent-assisted monitoring across product and infrastructure layers.

  • Analyse operational behaviour patterns across user personas and platform workloads to inform intelligent automation design and monitoring strategy evolution.

  • Define and govern Service Level Indicators and Objectives — using error budget data to drive engineering prioritisation and calibrate automation intervention thresholds rather than to reactively defend the committed SLA.

  • Define, instrument, and own platform reliability metrics — QoS, Uptime, MTTR, MTBF, and agent automation coverage — as leading indicators of system health and team maturity.

  • Synthesise and publish key metrics to stakeholders

  • Leverage deep AWS expertise and DevOps toolchain knowledge to design infrastructure automation that operates with the reliability and predictability of a software system.

  • Continuously analyse infrastructure and tooling spend — identifying waste, right-sizing opportunities, and cost anomalies through automated FinOps signal processing.

  • Build and operationalise cost governance frameworks where agent-driven policy enforcement, not periodic human review, is the primary control mechanism.

  • Design and implement AI-augmented observability solutions across the stack — integrating SumoLogic, CloudWatch, Grafana, Prometheus, and PagerDuty with agentic reasoning layers that interpret signal and act, not just alert.

  • Lead outage management with agentic support — automated problem detection, structured stakeholder communication, and agent-assisted resolution with clear human escalation protocols for high-severity events.

  • Own incident management and disaster recovery strategy — with automated runbook execution, agent-supervised recovery playbooks, and validated DR testing cadences.

  • Convert institutional knowledge into machine-executable runbooks and agent-accessible knowledge bases — ensuring operational intelligence is codified, versioned, and continuously improved

  • Lead operational improvement through continuous automation — replacing recurring manual processes with agent-driven workflows and measuring success by the reduction in human intervention per unit of platform activity.

  • Conduct rigorous incident postmortems and RCAs — using AI-assisted log correlation and timeline reconstruction to surface systemic gaps faster and drive durable corrective action.

  • Partner with development, QA, and architecture teams to embed reliability and automation requirements early in the SDLC — shifting reliability left rather than absorbing complexity at the production boundary.

  • Contribute to system design reviews with a reliability and automation lens — ensuring that agentic operability, observability hooks, and failure mode handling are first-class design considerations, not afterthoughts.

  • Provide senior technical leadership — setting the bar for agentic automation design, code quality in SRE tooling, and engineering rigour across the team.

  • Demonstrate Ownership and accountability.

Qualifications & Skills required

This is a critical service delivery role requiring experience with complex datacenter and cloud hosting environments. Pitney Bowes product hosting solutions leverage multiple technologies in complex data center and cloud environments that support multi-tiered high-availability applications. The role requires a talented self-directed and self-motivated individual with a strong work ethic and the following skills:

  • Graduate or Post-Graduate in Computer Science, Engineering, or a related discipline — or equivalent demonstrated depth through professional experience.

  • 10+ years of SRE or platform engineering experience, with a demonstrable shift in recent years toward automation-first and AI-augmented operations.

  • Excellent written and verbal communication skills — able to translate agent behaviour, reliability signals, and automation outcomes into clear engineering and executive narratives.

  • Strong background in SaaS product operations — with hands-on experience running multi-region, high-availability platforms at enterprise scale.

  • Strong experience with CI/CD tooling including Git and Argo — with the ability to extend pipelines with intelligent gates, automated validation, and agent-triggered rollback logic.

  • Strong experience with Docker, Kubernetes, microservices orchestration, and container lifecycle management — including automated remediation of cluster-level failure patterns.

  • Deep expertise in AWS services and cloud-native architectures — including event-driven automation, Lambda-based remediation, and cloud control plane integration for agentic workflows.

  • Solid grounding in networking fundamentals and firewall concepts — sufficient to design and debug connectivity for distributed cloud services and agentic tool integrations.

  • Familiarity with PaloAlto firewall is an advantage, particularly for candidates with GovCloud or regulated-environment experience.

  • Strong experience with Infrastructure as Code — Terraform, Ansible, CloudFormation — with a focus on policy-driven, drift-detecting, and self-correcting infrastructure patterns.

  • Strong Python proficiency — including building agentic tool integrations, LLM API wrappers, and operational automation scripts. Shell scripting for platform automation. PowerShell where Windows surfaces require it.

  • Strong hands-on experience across the observability stack — Prometheus, Grafana, SumoLogic, CloudWatch, PagerDuty, and OpsGenie — with the ability to extend these platforms with custom agentic reasoning and automated response layers.

  • Strong working knowledge of Linux — process management, networking, filesystem, and kernel-level troubleshooting. Windows familiarity where platform surfaces require it.

  • Strong analytical and systems debugging skills — able to reason through complex distributed failure modes and translate that understanding into automated detection and remediation logic.

  • Proficient with Jira, Confluence, and SharePoint — and comfortable integrating these platforms into agentic workflows for automated ticket creation, runbook retrieval, and knowledge surfacing.

  • Experienced working within Agile delivery models — contributing to sprint planning, backlog grooming, and cross-team engineering ceremonies as a senior technical voice.

  • Strong cross-functional collaboration capability — working fluidly with Product Engineering, Product Management, Client Success, and senior leadership to align reliability investments with business outcomes.

  • Hands-on experience designing and operating agentic systems — using frameworks such as LangGraph, AutoGen, or CrewAI to build multi-step, tool-calling agent workflows for operational use cases.

  • Practical experience with LLM APIs (OpenAI, Anthropic, or equivalent) including prompt engineering, tool/function calling, structured output design, and integrating AI reasoning into platform automation pipelines.

  • Ability to design human-in-the-loop escalation protocols for agentic systems — defining clear intervention thresholds, override mechanisms, and audit trails that keep autonomous operations safe and auditable.

  • Experience with agent observability and guardrails — instrumenting agent reasoning chains, detecting failure modes in autonomous workflows, and implementing safety boundaries that prevent runaway automation.

  • High personal accountability — for the reliability of systems you own, the quality of automation you build, and the engineering standards you set for others. 

About Pitney Bowes

Pitney Bowes (NYSE:PBI) is a global technology company providing commerce solutions that power billions of transactions. Clients around the world, including 90 percent of the Fortune 500, rely on the accuracy and precision delivered by Pitney Bowes solutions, analytics, and APIs in the areas of ecommerce fulfillment, shipping and returns; cross-border ecommerce; office mailing and shipping; presort services; and financing. For 100 years Pitney Bowes has been innovating and delivering technologies that remove the complexity of getting commerce transactions precisely right. For additional information visit Pitney Bowes at https://www.pitneybowes.com/in.

We will:


• Provide the will: opportunity to grow and develop your career
• Offer an inclusive environment that encourages diverse perspectives and ideas
• Deliver challenging and unique opportunities to contribute to the success of a transforming organization
• Offer comprehensive benefits globally (PB Benefits and Wellbeing Programs)

Pitney Bowes is an equal opportunity employer that values diversity and inclusiveness in the workplace.
All interested individuals must apply online.

Skills Required

  • Graduate or postgraduate degree in Computer Science, Engineering, or a related discipline, or equivalent professional experience
  • 10+ years of SRE or platform engineering experience
  • Experience with automation-first and AI-augmented operations
  • Strong SaaS product operations experience running multi-region, high-availability enterprise platforms
  • Strong CI/CD experience with Git and Argo
  • Experience with Docker, Kubernetes, microservices orchestration, and container lifecycle management
  • Deep expertise in AWS services and cloud-native architectures
  • Networking fundamentals and firewall concepts
  • Familiarity with Palo Alto firewalls
  • Strong Infrastructure as Code experience with Terraform, Ansible, and CloudFormation
  • Strong Python proficiency and experience with shell scripting; PowerShell for Windows environments
  • Hands-on observability experience with Prometheus, Grafana, SumoLogic, CloudWatch, PagerDuty, and OpsGenie
  • Strong Linux knowledge and Windows familiarity
  • Strong analytical and distributed-systems debugging skills
  • Proficiency with Jira, Confluence, and SharePoint
  • Experience working in Agile delivery models
  • Strong cross-functional collaboration and communication skills
  • Hands-on experience designing and operating agentic systems using LangGraph, AutoGen, or CrewAI
  • Practical experience with LLM APIs, prompt engineering, tool or function calling, structured outputs, and AI platform automation
  • Ability to design human-in-the-loop escalation protocols, override mechanisms, and audit trails
  • Experience with agent observability, autonomous workflow failure detection, and guardrails
  • High personal accountability for system reliability, automation quality, and engineering standards

Pitney Bowes Inc. Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Pitney Bowes Inc. and has not been reviewed or approved by Pitney Bowes Inc..

  • Healthcare Strength Health coverage includes medical, dental, and vision options plus mental health support, FSAs/HSAs, an Employee Assistance Program, and wellness offerings. Feedback suggests these features are a notable bright spot within total rewards.
  • Retirement Support Retirement offerings include a 401(k) with company match and a pension plan, alongside financial protections and an Employee Stock Purchase Plan. These elements pair with solid insurance to strengthen overall financial security.
  • Parental & Family Support Family supports include paid parental leave and adoption assistance, complemented by family medical leave. These programs align with broader leave options to support caregiving needs.

Pitney Bowes Inc. Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Stamford, CT
12,066 Employees
Year Founded: 1920

What We Do

Pitney Bowes (NYSE:PBI) is a global shipping and mailing company that provides technology, logistics, and financial services to more than 90 percent of the Fortune 500. Small business, retail, enterprise, and government clients around the world rely on Pitney Bowes to remove the complexity of sending mail and parcels. For additional information visit Pitney Bowes at www.pitneybowes.com.

Similar Jobs

Pitney Bowes Inc. Logo Pitney Bowes Inc.

Software Engineer

Information Technology • Logistics • Financial Services
In-Office
Pune, Mahārāshtra, IND
12066 Employees

Zocdoc Logo Zocdoc

Sales Development Representative

Healthtech • Information Technology • Software • Telehealth
Easy Apply
Hybrid
Pune, Mahārāshtra, IND
900 Employees

CSC Logo CSC

Associate Client Order Coordinator

Fintech • Legal Tech • Software • Financial Services • Cybersecurity • Data Privacy
Remote or Hybrid
2 Locations
8500 Employees

Zocdoc Logo Zocdoc

Operations Associate

Healthtech • Information Technology • Software • Telehealth
Easy Apply
Hybrid
Pune, Mahārāshtra, IND
900 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account