The Custom Agent Service handles the full agent lifecycle: an admin configures an agent with instructions, knowledge articles, and actions through a UI; a customer ticket triggers execution; the agent plans, acts through real APIs via the Integration Action Platform, and resolves the issue. The backend is Python, the infrastructure is Kubernetes on AWS, and the agent architectures range from single-pass ReAct loops to our iterative multi-plan executor.
We are onboarding our first internal and external EAP customers, so the work ships to real accounts with real tickets.
What we need help withAgent execution core. The planning loop, tool dispatch, memory integration, and error recovery that make up the main execution path. You will work directly on the code that decides what the agent does next and make it faster, more reliable, and more capable. This includes integrating new architectures into the production path and hardening them for real traffic.
Knowledge retrieval. Agents retrieve and reason over customer knowledge bases at runtime. The retrieval pipeline (embedding, reranking, context assembly) needs to balance answer quality against latency and token cost, across thousands of heterogeneous knowledge bases per deployment.
Actions and connectors. Agents call Zendesk APIs, third-party connectors (Shopify, Salesforce, etc.), custom actions configured by admins, and increasingly other agents via A2A. The execution layer needs reliable retries, timeouts, schema validation, and graceful degradation when connectors fail mid-execution. You would also self-service new connector integrations through the Connector SDK.
Production instrumentation for model training. Every agent execution generates a trajectory (reasoning steps, tool calls, outcomes, user feedback). We are building toward training domain-specialized models, and that requires clean production data. You would instrument the execution pipeline to capture implicit reward signals (resolution success, escalation patterns, user satisfaction) that feed into the ML team's training pipeline.
Security and compliance. PII filtering, audit logging, action versioning, and governance patterns that keep agents within admin-configured bounds. You would work directly with Product Security on security review items as the platform scales.
What we are looking for- 5+ years of backend engineering with strong Python skills. You have shipped production systems, not just models. You understand the difference between getting an agent to work locally and running it across 100,000 accounts.
- Comfortable across the full agent stack: LLM APIs, prompt engineering, tool calling, memory management, evaluation. You can build an agent loop from scratch, and you know when a framework helps vs. when it gets in the way.
- You think about what happens when the model returns garbage, the connector times out, and the customer is waiting. You build for the failure case, not just the happy path.
- You ship working code, review PRs carefully, and communicate clearly about what is done, what is blocked, and what is at risk.
- Languages: Python (primary), some Go for platform services
- Agent Frameworks: Custom iterative architectures, ReAct, with integration points to open-source tooling
- Infrastructure: Kubernetes, Spinnaker, AWS (ECS, S3, ElastiCache)
- Data: Postgres, ElastiCache/Redis (vector + KV), Kafka
- Evaluation: Braintrust (experiment tracking, scoring, CI/CD integration)
- Protocols: MCP, REST, gRPC
Why Zendesk for this work
Zendesk has 100,000+ customers, billions of support interactions, and a live product surface where agents are already resolving tickets. The feedback loop from an agent action to a measurable customer outcome is minutes, not months. We are hiring 2+ engineers at each of these levels.
The intelligent heart of customer experience
Zendesk software was built to bring a sense of calm to the chaotic world of customer service. Today we power billions of conversations with brands you know and love.
Zendesk believes in offering our people a fulfilling and inclusive experience. Our hybrid way of working, enables us to purposefully come together in person, at one of our many Zendesk offices around the world, to connect, collaborate and learn whilst also giving our people the flexibility to work remotely for part of the week.
As part of our commitment to fairness and transparency, we inform all applicants that artificial intelligence (AI) or automated decision systems may be used to screen or evaluate applications for this position, in accordance with Company guidelines and applicable law.
Zendesk is an equal opportunity employer, and we’re proud of our ongoing efforts to foster global diversity, equity, & inclusion in the workplace. Individuals seeking employment and employees at Zendesk are considered without regard to race, color, religion, national origin, age, sex, gender, gender identity, gender expression, sexual orientation, marital status, medical condition, ancestry, disability, military or veteran status, or any other characteristic protected by applicable law. We are an AA/EEO/Veterans/Disabled employer. If you are based in the United States and would like more information about your EEO rights under the law, please click here.
Zendesk endeavors to make reasonable accommodations for applicants with disabilities and disabled veterans pursuant to applicable federal and state law. If you are an individual with a disability and require a reasonable accommodation to submit this application, complete any pre-employment testing, or otherwise participate in the employee selection process, please send an e-mail to [email protected] with your specific accommodation request.
Skills Required
- 5+ years of backend engineering experience with strong Python skills
- Proven experience shipping production systems (not just models) at scale
- Experience building agent stacks: LLM APIs, prompt engineering, tool calling, memory management, evaluation
- Designing and hardening agent execution: planning loops, tool dispatch, error recovery, and reliability
- Building knowledge retrieval pipelines (embeddings, reranking, context assembly) balancing latency and cost
- Integrating connectors and APIs with retries, timeouts, schema validation, and graceful degradation
- Instrumenting production for model training: capture trajectories, implicit reward signals, and evaluation metrics
- Security and compliance experience: PII filtering, audit logging, action versioning, governance
- Familiarity with Kubernetes and AWS infrastructure (ECS, S3, ElastiCache)
- Experience with Postgres, Redis/ElastiCache (vector + KV), Kafka, and protocols like REST/gRPC
- Familiarity with Go for platform services, Spinnaker, and connector ecosystems (Zendesk, Shopify, Salesforce)
- Ability to ship working code, perform careful PR reviews, and communicate status and risks clearly
Zendesk Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Zendesk and has not been reviewed or approved by Zendesk.
-
Fair & Transparent Compensation — The company states a commitment to publishing base pay ranges and advancing pay equity, helping employees gauge fairness. Public messaging on pay equity and transparency signals structured, consistent compensation practices.
-
Leave & Time Off Breadth — Time away programs include flexible PTO, dedicated well‑being days, emergency time off, and pregnancy loss leave. Parental leave is described as generous, and travel support exists for reproductive care where access is restricted.
-
Healthcare Strength — Benefits language highlights comprehensive medical, dental/vision, mental health access, and an employee assistance program. These offerings are positioned as part of holistic wellbeing support across regions.
Zendesk Insights
What We Do
Zendesk software was built to bring a sense of calm to the chaotic world of customer service. Today we power billions of conversations with brands you know and love. We advocate for digital first customer experiences— and we stick with it in our workplace. Over 5,000 employees worldwide are collaborating from kitchen tables, home offices, co-working spaces, and Zendesk workspaces to make one team.
Why Work With Us
We know one desk doesn’t fit all. At Zendesk, we prioritize remote work because we believe great work happens anywhere. Digital first is more than where we work though. We give our employees flexibility and choice in both where and how they work while also trusting them to be a team player.
Gallery









