About UOB
United Overseas Bank Limited (UOB) is a leading bank in ASEAN with a global network in Southeast Asia, Asia Pacific, Europe and North America. Operating through our head office in Singapore and banking subsidiaries in China, Indonesia, Malaysia, Thailand and Vietnam, we have a global network of about 430 branches and offices in 19 markets. At the heart of UOB is our culture, shaped by the UOB Way and anchored on our four values – Honourable, Enterprising, United and Committed. For more than 90 years, these values have guided how we do right by our customers, collaborate with one another and create long-term value for the communities we operate in. As One Bank, we are committed to helping our colleagues build sustainable careers grounded in purpose, supported by strong values, and enriched with meaningful opportunities to grow.
Job Description
- Own L3 incident response for GenAI use cases in production. Triage, deep debug, fix or work with the vendor to resolve defects and improve reliability
- Review, reproduce and root cause complex bugs across LangChain flows, Elasticsearch queries, Redis caches and safety layers such as Guardrails AI and LlamaGuard
- Implement safe changes and hotfixes. Build, test and deploy code and config updates through Jenkins with proper approvals and rollbacks
- Partner with business users to understand issue impact and edge cases. Translate vague problem reports into actionable steps and clear acceptance criteria
- Harden use case pipelines. Add input validation, timeouts, retries, circuit breakers and fallback strategies to reduce customer facing impact
- Improve RAG quality for stability. Tune chunking, retrieval parameters, query rewriting and caching to reduce latency and errors
- Maintain and evolve runbooks, playbooks and knowledge articles. Keep them current and prove they work with regular game days
- Drive observability for the use cases. Define golden signals, add OpenTelemetry traces, wire up Prometheus metrics and curate Grafana dashboards
- Manage production hygiene. Track error budgets, SLOs and SLAs. Push for defect burn down and change quality
- Coordinate with platform and infra teams when incidents touch OpenShift, GPUs, vLLM or NVIDIA Enterprise AI services
- Perform safe data operations. Validate indices, manage Redis eviction strategies and handle backfills and reindex tasks with minimal risk
- Champion secure by default practices. Enforce secrets handling, PII redaction and prompt or output guardrails
- Lead post incident reviews with blameless RCAs and concrete corrective actions. Close the loop with business stakeholders
- Proactively surface reliability risks and propose code or architecture changes that remove recurring failure modes
Job Requirement
- 3 to 4 years of software engineering or SRE with at least one year supporting AI or data intensive services in production
- Strong Python. Comfortable reading and fixing LangChain code, FastAPI or Flask services, async patterns and task queues
- Hands on with LangChain in real projects. Chains, tools, retrievers, memory and callback handlers
- Elasticsearch proficiency. Query DSL, relevance tuning, index lifecycle, scaling, snapshots and Kibana for triage
- Redis proficiency. Caching patterns, pub or sub, streams, eviction, persistence and troubleshooting latency or timeouts
- Safety and guardrails experience. Guardrails AI configuration, schema and validator design, LlamaGuard or similar classifiers in the loo
- CI or CD with Jenkins. Declarative pipelines, approvals, artifacts, environment promotion and rollback strategy
- Observability depth. Prometheus metrics, Grafana dashboards, OpenTelemetry traces and logs.
- Able to decide which signals matter and set actionable alerts
- Solid Git. Branching, pull requests, code reviews and release tagging. Comfortable with feature flags and canary or blue green patterns
- Working knowledge of Kubernetes and OpenShift as a strong plus. Debugging pods, logs, events and basic resource tuning
- Familiarity with LLM serving stacks such as vLLM or NVIDIA Enterprise AI is a plus. Know how to read model server logs, timeouts and token throughput
- Data handling discipline. Understanding of PII, masking, prompt redaction and safe logging
- Clear communicator who can write crisp incident updates, RCAs and user facing notes
Additional Requirements
Be a Part of the UOB Family
UOB is an equal opportunity employer. UOB does not discriminate on the basis of a candidate's age, race, gender, color, religion, sexual orientation, physical or mental disability, or other non-merit factors. All employment decisions at UOB are based on business needs, job requirements and qualifications. If you require any assistance or accommodations to be made for the recruitment process, please inform us when you submit your online application.
Apply now and make a Difference
Skills Required
- 3 to 4 years of software engineering or SRE experience
- At least 1 year supporting AI or data-intensive services in production
- Strong Python skills
- Experience with LangChain chains, tools, retrievers, memory, and callback handlers
- Experience with FastAPI or Flask services, asynchronous patterns, and task queues
- Elasticsearch proficiency, including Query DSL, relevance tuning, index lifecycle, scaling, snapshots, and Kibana
- Redis proficiency, including caching, pub/sub, streams, eviction, persistence, and latency troubleshooting
- Experience configuring safety guardrails, schema validators, Guardrails AI, LlamaGuard, or similar classifiers
- Jenkins CI/CD experience, including declarative pipelines, approvals, artifacts, environment promotion, and rollbacks
- Observability experience with Prometheus, Grafana, OpenTelemetry traces, and logs
- Ability to define actionable monitoring signals and alerts
- Git experience with branching, pull requests, code reviews, and release tagging
- Experience with feature flags and canary or blue-green deployment patterns
- Working knowledge of Kubernetes and OpenShift
- Familiarity with vLLM or NVIDIA Enterprise AI model-serving stacks
- Understanding of PII handling, masking, prompt redaction, and safe logging
- Clear written and verbal communication, including incident updates, RCAs, and user-facing notes
What We Do
We’re here to do Right By You. At UOB, we aspire to build a better future for the people and businesses in the region. Through our extensive network and suite of capabilities, we offer financial solutions to the people and businesses within, and connecting with ASEAN. We create solutions tailored to your unique needs through data and relationship-led insights. Our comprehensive regional network and one-bank approach connects your business to new opportunities in ASEAN. We help businesses to advance responsibly and guide personal wealth to grow sustainably. We foster inclusiveness and environmental well-being for stronger societies. This is how we stay committed to forging a sustainable future for generations to come. Note: For the terms of use of our LinkedIn channel, please visit: https://go.uob.com/socialmedia








