Manager, GenAI L3 Support Engineer, Group Technology & Ops

Posted 3 Days Ago
Be an Early Applicant
Hiring Remotely in Eastern Region, East, CMR
Remote
Mid level
Financial Services
The Role
Own L3 production incident response for GenAI applications, including triage, debugging, hotfixes, root-cause analysis, and reliability improvements. Harden RAG pipelines, improve observability, manage Elasticsearch and Redis operations, maintain runbooks, and coordinate with infrastructure and platform teams. Implement secure data handling, guardrails, deployment controls, and resilient production patterns. Lead post-incident reviews and drive corrective actions with business stakeholders.
Summary Generated by Built In
Company: 1011 United Overseas Bank Ltd

About UOB

United Overseas Bank Limited (UOB) is a leading bank in ASEAN with a global network in Southeast Asia, Asia Pacific, Europe and North America. Operating through our head office in Singapore and banking subsidiaries in China, Indonesia, Malaysia, Thailand and Vietnam, we have a global network of about 430 branches and offices in 19 markets. At the heart of UOB is our culture, shaped by the UOB Way and anchored on our four values – Honourable, Enterprising, United and Committed. For more than 90 years, these values have guided how we do right by our customers, collaborate with one another and create long-term value for the communities we operate in. As One Bank, we are committed to helping our colleagues build sustainable careers grounded in purpose, supported by strong values, and enriched with meaningful opportunities to grow.


Job Description

  • Own L3 incident response for GenAI use cases in production. Triage, deep debug, fix or work with the vendor to resolve defects and improve reliability
  • Review, reproduce and root cause complex bugs across LangChain flows, Elasticsearch queries, Redis caches and safety layers such as Guardrails AI and LlamaGuard
  • Implement safe changes and hotfixes. Build, test and deploy code and config updates through Jenkins with proper approvals and rollbacks
  • Partner with business users to understand issue impact and edge cases. Translate vague problem reports into actionable steps and clear acceptance criteria
  • Harden use case pipelines. Add input validation, timeouts, retries, circuit breakers and fallback strategies to reduce customer facing impact
  • Improve RAG quality for stability. Tune chunking, retrieval parameters, query rewriting and caching to reduce latency and errors
  • Maintain and evolve runbooks, playbooks and knowledge articles. Keep them current and prove they work with regular game days
  • Drive observability for the use cases. Define golden signals, add OpenTelemetry traces, wire up Prometheus metrics and curate Grafana dashboards
  • Manage production hygiene. Track error budgets, SLOs and SLAs. Push for defect burn down and change quality
  • Coordinate with platform and infra teams when incidents touch OpenShift, GPUs, vLLM or NVIDIA Enterprise AI services
  • Perform safe data operations. Validate indices, manage Redis eviction strategies and handle backfills and reindex tasks with minimal risk
  • Champion secure by default practices. Enforce secrets handling, PII redaction and prompt or output guardrails
  • Lead post incident reviews with blameless RCAs and concrete corrective actions. Close the loop with business stakeholders
  • Proactively surface reliability risks and propose code or architecture changes that remove recurring failure modes

Job Requirement

  • 3 to 4 years of software engineering or SRE with at least one year supporting AI or data intensive services in production
  • Strong Python. Comfortable reading and fixing LangChain code, FastAPI or Flask services, async patterns and task queues
  • Hands on with LangChain in real projects. Chains, tools, retrievers, memory and callback handlers
  • Elasticsearch proficiency. Query DSL, relevance tuning, index lifecycle, scaling, snapshots and Kibana for triage
  • Redis proficiency. Caching patterns, pub or sub, streams, eviction, persistence and troubleshooting latency or timeouts
  • Safety and guardrails experience. Guardrails AI configuration, schema and validator design, LlamaGuard or similar classifiers in the loo
  • CI or CD with Jenkins. Declarative pipelines, approvals, artifacts, environment promotion and rollback strategy
  • Observability depth. Prometheus metrics, Grafana dashboards, OpenTelemetry traces and logs.
  • Able to decide which signals matter and set actionable alerts
  • Solid Git. Branching, pull requests, code reviews and release tagging. Comfortable with feature flags and canary or blue green patterns
  • Working knowledge of Kubernetes and OpenShift as a strong plus. Debugging pods, logs, events and basic resource tuning
  • Familiarity with LLM serving stacks such as vLLM or NVIDIA Enterprise AI is a plus. Know how to read model server logs, timeouts and token throughput
  • Data handling discipline. Understanding of PII, masking, prompt redaction and safe logging
  • Clear communicator who can write crisp incident updates, RCAs and user facing notes

Additional Requirements

Be a Part of the UOB Family

UOB is an equal opportunity employer. UOB does not discriminate on the basis of a candidate's age, race, gender, color, religion, sexual orientation, physical or mental disability, or other non-merit factors. All employment decisions at UOB are based on business needs, job requirements and qualifications. If you require any assistance or accommodations to be made for the recruitment process, please inform us when you submit your online application.

Apply now and make a Difference

Skills Required

  • 3 to 4 years of software engineering or SRE experience
  • At least 1 year supporting AI or data-intensive services in production
  • Strong Python skills
  • Experience with LangChain chains, tools, retrievers, memory, and callback handlers
  • Experience with FastAPI or Flask services, asynchronous patterns, and task queues
  • Elasticsearch proficiency, including Query DSL, relevance tuning, index lifecycle, scaling, snapshots, and Kibana
  • Redis proficiency, including caching, pub/sub, streams, eviction, persistence, and latency troubleshooting
  • Experience configuring safety guardrails, schema validators, Guardrails AI, LlamaGuard, or similar classifiers
  • Jenkins CI/CD experience, including declarative pipelines, approvals, artifacts, environment promotion, and rollbacks
  • Observability experience with Prometheus, Grafana, OpenTelemetry traces, and logs
  • Ability to define actionable monitoring signals and alerts
  • Git experience with branching, pull requests, code reviews, and release tagging
  • Experience with feature flags and canary or blue-green deployment patterns
  • Working knowledge of Kubernetes and OpenShift
  • Familiarity with vLLM or NVIDIA Enterprise AI model-serving stacks
  • Understanding of PII handling, masking, prompt redaction, and safe logging
  • Clear written and verbal communication, including incident updates, RCAs, and user-facing notes
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Singapore
25,000 Employees
Year Founded: 1935

What We Do

We’re here to do Right By You. At UOB, we aspire to build a better future for the people and businesses in the region. Through our extensive network and suite of capabilities, we offer financial solutions to the people and businesses within, and connecting with ASEAN. We create solutions tailored to your unique needs through data and relationship-led insights. Our comprehensive regional network and one-bank approach connects your business to new opportunities in ASEAN. We help businesses to advance responsibly and guide personal wealth to grow sustainably. We foster inclusiveness and environmental well-being for stronger societies. This is how we stay committed to forging a sustainable future for generations to come. Note: For the terms of use of our LinkedIn channel, please visit: https://go.uob.com/socialmedia

Similar Jobs

Falcon Funded Logo Falcon Funded

UGC Creator - Cameroon

Fintech • Financial Services
Remote
CM
125 Employees

Synchrony Logo Synchrony

Architect

Fintech • Financial Services
In-Office or Remote
26 Locations
10001 Employees

Synchrony Logo Synchrony

Data Architect

Fintech • Financial Services
In-Office or Remote
5 Locations
10001 Employees

Similar Companies Hiring

Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account