Role Overview
We are seeking a Vice President – AI Safety Platforms to build and lead our enterprise AI safety engineering initiatives. As generative AI in financial services evolves from simple prompt-response workflows to autonomous agentic systems that execute multi-step plans, call APIs, and interact directly with internal systems, establishing robust safety mechanisms and standardized evaluation protocols is essential.
In this role, you will recruit and lead dedicated engineering pods focused on developing a unified company-wide agentic evaluation framework, real-time LLM guardrail services, and automated governance controls. As the senior technical authority for AI safety, you will collaborate closely with core AI platform teams, risk control functions, and business units to drive necessary enhancements to the core AI platform (such as telemetry hooks, API capabilities, execution sandboxes, and data logging infrastructure) to ensure all enterprise AI deployments operate safely, verifiably, and in compliance with institutional standards.
Key Responsibilities
1. Unified Agentic Evaluation Framework
- Company-Wide Architecture: Design, build, and deploy a single, company-wide agentic evaluation framework that standardizes how teams across all business lines benchmark, test, and measure AI agent performance prior to production deployment.
- Trajectory & Multi-Step Reasoning Assessment: Implement evaluation methodologies that score autonomous planning quality, tool-calling precision, multi-turn state retention, trajectory efficiency, and error-recovery behaviors.
- Continuous Monitoring & Production Drift: Integrate automated evaluation pipelines into runtime environments to continuously audit agent execution traces, detecting reasoning drift, tool failure modes, and unexpected trajectory shifts in production.
- Domain-Specific Benchmarking: Establish standardized test suites and synthetic evaluation benchmarks tailored to complex financial workflows, such as automated research, risk assessment, and operational task execution.
2. LLM Guardrails Infrastructure & Real-Time Controls
- Low-Latency Guardrail Engine: Architect and scale enterprise guardrail microservices that inspect prompt inputs, retrieved context, and model outputs in real time to prevent data leakage, policy violations, and unvalidated execution.
- Tool-Use & Action Control: Implement runtime policy gateways that inspect and authorize tool calls before execution, ensuring agents operate within authorized data boundaries and action scopes.
- Human-in-the-Loop (HITL) Triggers: Build configurable escalation workflows and approval gates that automatically pause execution for high-risk operations (e.g., money movement, client record modifications, or external communications) until human authorization is granted.
3. Core AI Platform Enhancements & Governance Integration
- Drive Platform Enhancements: Partner directly with the core AI Platform team to drive the implementation of safety APIs, telemetry hooks, developer SDKs, and MLOps/LLMOps pipeline integrations.
- Auditability & Execution Telemetry: Define and enforce technical standards for immutable audit logging, execution tracing (e.g., OpenTelemetry standards), and principal identity propagation across all agentic workflows.
- Regulatory & Model Risk Alignment: Translate model risk management standards (e.g., SR 11-7 / SR 26-2 guidance, FINRA supervision requirements) into automated engineering safeguards and policy checks.
4. Engineering Leadership & Strategic Oversight
- Team Building & Mentorship: Hire, develop, and mentor high-performing engineering teams specializing in applied machine learning, AI safety, and enterprise platform engineering.
- Strategic Roadmap: Own the technical roadmap for enterprise AI safety infrastructure, setting clear milestones for evaluation framework adoption, runtime latency optimization, and governance automation.
- Stakeholder Collaboration: Articulate technical risk profiles, evaluation metrics, and safety architecture to risk committees, model validation teams, and executive leadership.
Key Qualifications
Basic Qualifications
- Role Level: Vice President experience (or equivalent senior engineering leadership) in financial services or large-scale enterprise software environments.
- Education: Bachelor’s or Master’s degree in Computer Science, Artificial Intelligence, Systems Engineering, or a related quantitative field.
- Engineering Leadership: 4+ years leading applied ML or software engineering teams in building platform infrastructure or microservices.
- Software Engineering Depth: 8+ years of hands-on software development experience (Python, Go, Java, or C++) building microservices, high-throughput APIs, or enterprise platform services.
- AI & Agentic Expertise: Technical fluency with Large Language Models (LLMs), RAG systems, function calling / tool integration, and agentic execution paradigms (e.g., LangChain, AutoGen, CrewAI, MCP server architectures).
Preferred Experience & Technical Skills
- Agentic Evaluation: Direct experience building agent evaluation frameworks and metrics (e.g., LLM-as-a-Judge, G-Eval, trajectory trace evaluation, task completion scoring).
- Guardrail Frameworks: Hands-on experience integrating low-latency guardrail tools and runtime filters (e.g., NeMo Guardrails, Guardrails AI, Llama Guard).
- AI Observability & Tracing: Experience with LLM and agent tracing tools (e.g., LangSmith, OpenTelemetry, Phoenix, MLflow) and structured audit logging infrastructure.
- Platform Engineering Alignment: Proven ability to partner across teams and drive key governance capabilities into core shared platforms.
Salary Range
The expected base salary for this New York, NY, United States-based position is $130000-$250000. In addition, you may be eligible for a discretionary bonus if you are an active employee as of fiscal year-end.
Benefits
Goldman Sachs is committed to providing our people with valuable and competitive benefits and wellness offerings, as it is a core part of providing a strong overall employee experience. A summary of these offerings, which are generally available
to active, non-temporary, full-time and part-time US employees who work at least 20 hours per week, can be found here.
Skills Required
- Vice President experience or equivalent senior engineering leadership in financial services or large-scale enterprise software environments
- Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Systems Engineering, or a related quantitative field
- 4+ years leading applied machine learning or software engineering teams building platform infrastructure or microservices
- 8+ years of hands-on software development experience using Python, Go, Java, or C++ to build microservices, high-throughput APIs, or enterprise platform services
- Technical fluency with LLMs, RAG systems, function calling, tool integration, and agentic execution paradigms
- Direct experience building agent evaluation frameworks and metrics, including LLM-as-a-Judge, G-Eval, trajectory trace evaluation, or task completion scoring
- Hands-on experience integrating low-latency guardrail tools and runtime filters such as NeMo Guardrails, Guardrails AI, or Llama Guard
- Experience with LLM and agent tracing tools such as LangSmith, OpenTelemetry, Phoenix, or MLflow, plus structured audit logging infrastructure
- Ability to partner across teams and drive governance capabilities into core shared platforms
Goldman Sachs Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Goldman Sachs and has not been reviewed or approved by Goldman Sachs.
-
Healthcare Strength — Coverage includes medical, dental, vision, disability, life and accident insurance, with multiple plan options and most premiums subsidized; coverage often starts on day one. Wellness resources, on-site health centers in some locations, and EAP access reinforce the depth of health support.
-
Parental & Family Support — Family care includes on-site childcare in some offices, expectant parent resources, and transitional programs for returning parents. Feedback suggests parental leave is very generous, with reports of around 20 weeks paid leave and stipends for adoption, surrogacy, and fertility-related services.
-
Retirement Support — The firm provides a 401(k) plan with employer matching contributions and broad financial education to help employees plan for retirement. Resources also support saving for education and preparing for unexpected events.
Goldman Sachs Insights
What We Do
At Goldman Sachs, we believe progress is everyone’s business. That’s why we commit our people, capital and ideas to help our clients, shareholders and the communities we serve to grow. Founded in 1869, Goldman Sachs is a leading global investment banking, securities and investment management firm. Headquartered in New York, we maintain offices in all major financial centers around the world. More about our company can be found at www.goldmansachs.com
.png)
.png)







