We are looking for an experienced Senior AI Infrastructure Engineer to join our team. This role owns the technical evaluation and end-to-end execution of AI infrastructure deployments — from data center due diligence and solution review, through cross- functional delivery coordination, to ongoing operations of production AI environments. You will act as a key technical owner bridging internal stakeholders and external partners, ensuring infrastructure is delivered on time, to spec, and operated reliably at scale.Key Responsibilities
• Conduct on-site data center assessments to evaluate whether candidate facilities meet AI infrastructure requirements.
• Review and validate AI infrastructure solution designs, identifying technical risks, gaps, and cost/performance trade-offs before sign-off.
• Coordinate with external partners, driving timelines, resolving technical issues, and ensuring deliverables meet internal requirements.
• Partner with internal business and engineering teams to translate requirements into deliverable technical specifications.
• Participate in and eventually take ownership of day-2 operations for live AI
environments — monitoring, incident response, capacity/health checks, firmware and lifecycle management, coordinating hardware maintenance as needed.
• Build and improve operational standards, runbooks, and documentation for
infrastructure delivery and operations to enable consistent execution across regions.
• Provide technical guidance and mentorship to junior engineers/interns on hardware diagnostics, cluster tooling, and best practices.
• Track and report on project status and risks to management.
Who We Look ForRequirements• Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
• 5+ years of experience in data center infrastructure, infrastructure deployment, or infrastructure operations.
• Solid understanding of high-performance compute server hardware, high-speed
networking (e.g., InfiniBand/RoCE), storage systems, and data center power & cooling fundamentals.
• Proven experience evaluating data center facilities and reviewing technical proposals for compute infrastructure.
• Strong track record coordinating complex, multi-party technical projects to closure, including working effectively with external partners and cross-regional teams.
• Hands-on experience with cluster orchestration and management tools (e.g., Slurm, Kubernetes) and Linux system administration.
• Proficiency in scripting/automation (Python, Bash; Ansible/Terraform a plus).
• Familiarity with observability/monitoring stacks (Prometheus, Grafana, ELK) and DCIM tooling.
• Bilingual proficiency in English and Mandarin highly preferred
• Strong ownership mentality, able to independently drive projects with minimal
supervision.
Preferred Qualifications• Experience with high-performance storage / parallel file systems (e.g., Lustre,
GPFS/Spectrum Scale, WekaFS, VAST, Ceph) in AI infrastructure environments.
• Experience deploying/operating large-scale AI compute environments in a hyperscale or cloud environment.
• Data center or infrastructure certifications.
Location State(s)
US-California-Palo AltoThe expected base pay range for this position in the location(s) listed above is $124,800.00 to $283,800.00 per year. Actual pay may vary depending on job-related knowledge, skills, and experience. Employees hired for this position may be eligible for a sign on payment, relocation package, and restricted stock units, which will be evaluated on a case-by-case basis. Subject to the terms and conditions of the plans in effect, hired applicants are also eligible for medical, dental, vision, life and disability benefits, and participation in the Company’s 401(k) plan. The Employee is also eligible for up to 15 to 25 days of vacation per year (depending on the employee’s tenure), up to 13 days of holidays throughout the calendar year, and up to 10 days of paid sick leave per year. Your benefits may be adjusted to reflect your location, employment status, duration of employment with the company, and position level. Benefits may also be pro-rated for those who start working during the calendar year.Equal Employment Opportunity at TencentAs an equal opportunity employer, we firmly believe that diverse voices fuel our innovation and allow us to better serve our users and the community. We foster an environment where every employee of Tencent feels supported and inspired to achieve individual and common goals.
Skills Required
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field
- 5+ years of experience in data center infrastructure, infrastructure deployment, or infrastructure operations
- Understanding of high-performance compute server hardware, high-speed networking such as InfiniBand or RoCE, storage systems, and data center power and cooling fundamentals
- Experience evaluating data center facilities and reviewing technical proposals for compute infrastructure
- Experience coordinating complex, multi-party technical projects through completion, including external partners and cross-regional teams
- Hands-on experience with cluster orchestration and management tools such as Slurm or Kubernetes
- Linux system administration experience
- Proficiency in Python and Bash scripting or automation
- Familiarity with observability and monitoring stacks such as Prometheus, Grafana, and ELK
- Familiarity with DCIM tooling
- Strong ownership and ability to independently drive projects with minimal supervision
- Bilingual proficiency in English and Mandarin
- Experience with high-performance storage or parallel file systems such as Lustre, GPFS/Spectrum Scale, WekaFS, VAST, or Ceph
- Experience deploying or operating large-scale AI compute environments in a hyperscale or cloud environment
- Data center or infrastructure certifications
- Experience with Ansible or Terraform
What We Do
Tencent is a world-leading internet and technology company founded in 1998 in Shenzhen, China. It develops products and services spanning social platforms and messaging, online games, digital media, payments and financial technology, cloud and advertising solutions, utility software, and enterprise digital transformation. Its mission is to use technology for good and improve the quality of life of people around the world.








