AI Systems Performance Specialist

Sorry, this job was removed at 08:22 p.m. (UTC) on Monday, Aug 31, 2026
Be an Early Applicant
Chapel Hill, NC, USA
In-Office
130K-180K Annually
Expert/Leader
Artificial Intelligence • Information Technology • Software • Consulting
The Role
Optimize enterprise-scale AI training and inference systems for throughput, latency, scalability, reliability, and cost efficiency. Responsibilities include GPU and multi-GPU optimization, distributed training, LLM inference, profiling, model optimization, benchmarking, monitoring, cloud and FinOps optimization, and evaluating emerging AI hardware and frameworks. The role collaborates across AI, ML, platform, and infrastructure teams while providing technical leadership and mentoring.
Summary Generated by Built In

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.

Job Title

AI Systems Performance Specialist

Location: 100% Remote (Continental United States)
Position Type: Full-time, Direct W2
Salary Range: $130,000–$180,000 Annually
Experience Required: 10+ Years

Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.

Job Summary

Bright Vision Technologies is seeking a highly experienced AI Systems Performance Specialist with 10+ years of experience in AI infrastructure, machine learning systems, High-Performance Computing (HPC), and performance engineering. The ideal candidate will optimize AI training and inference workloads for maximum performance, scalability, reliability, and cost efficiency. This role requires deep expertise in GPU optimization, distributed training, Large Language Model (LLM) inference, Python, C++, CUDA, and production AI systems, along with the ability to lead performance optimization initiatives across enterprise-scale AI platforms.

Key Responsibilities
  • Optimize AI training and inference pipelines for maximum throughput, low latency, scalability, and infrastructure efficiency.
  • Analyze and improve GPU utilization, memory management, kernel execution, and multi-GPU performance across production AI workloads.
  • Design and implement optimization techniques including quantization, pruning, mixed precision, batching, caching, speculative decoding, and model parallelism.
  • Profile AI applications using industry-standard performance analysis tools and identify bottlenecks across compute, memory, networking, and storage.
  • Optimize distributed training and inference using NCCL, DeepSpeed, PyTorch Distributed, Ray, MPI, or similar distributed computing frameworks.
  • Collaborate with AI researchers, ML engineers, platform engineers, and infrastructure teams to improve model performance and production reliability.
  • Build automated benchmarking frameworks, performance dashboards, monitoring solutions, and regression testing pipelines.
  • Evaluate emerging AI hardware, GPU architectures, inference frameworks, and optimization technologies to improve enterprise AI capabilities.
  • Drive AI infrastructure cost optimization through efficient resource utilization, cloud optimization, and FinOps best practices.
  • Mentor engineering teams and provide technical leadership on AI systems architecture, GPU optimization, and performance engineering.
Required Qualifications
  • Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, or a related technical discipline.
  • 10+ years of professional experience in performance engineering, AI infrastructure, machine learning systems, High-Performance Computing (HPC), or distributed computing.
  • Expert-level programming skills in Python and C++.
  • Extensive experience optimizing GPU-accelerated AI workloads using CUDA, distributed training frameworks, and modern deep learning libraries.
  • Strong knowledge of Large Language Models (LLMs), deep learning frameworks, model serving, and production AI inference.
  • Hands-on experience with profiling tools such as NVIDIA Nsight Systems, Nsight Compute, PyTorch Profiler, TensorBoard, or similar performance analysis tools.
  • Experience deploying and optimizing AI workloads on AWS, Microsoft Azure, or Google Cloud Platform (GCP).
  • Strong understanding of distributed systems, networking, storage optimization, and AI infrastructure architecture.
  • Excellent analytical, troubleshooting, communication, and technical leadership skills.
Preferred Qualifications
  • Experience optimizing production-scale LLM inference and serving large foundation models.
  • Hands-on experience with vLLM, TensorRT-LLM, DeepSpeed, Triton Inference Server, CUTLASS, FasterTransformer, or similar AI optimization frameworks.
  • Knowledge of model compression, KV cache optimization, speculative decoding, and advanced inference optimization techniques.
  • Experience implementing FinOps strategies for AI infrastructure cost optimization and resource management.
  • Contributions to AI systems research, open-source AI infrastructure projects, patents, or technical publications.
  • Familiarity with emerging AI accelerator technologies, including AMD ROCm, Intel oneAPI, or custom AI hardware.

Interested in this opportunity? Apply today for immediate consideration!
Email your updated resume: [email protected]
Call or Text: (908) 505-3545
Learn more: www.bvteck.com

Bright Vision Technologies is an Equal Opportunity Employer.

Skills Required

  • Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, or a related technical discipline
  • 10+ years of professional experience in performance engineering, AI infrastructure, machine learning systems, High-Performance Computing, or distributed computing
  • Expert-level programming skills in Python and C++
  • Extensive experience optimizing GPU-accelerated AI workloads using CUDA, distributed training frameworks, and modern deep learning libraries
  • Strong knowledge of Large Language Models, deep learning frameworks, model serving, and production AI inference
  • Hands-on experience with NVIDIA Nsight Systems, Nsight Compute, PyTorch Profiler, TensorBoard, or similar performance analysis tools
  • Experience deploying and optimizing AI workloads on AWS, Microsoft Azure, or Google Cloud Platform
  • Strong understanding of distributed systems, networking, storage optimization, and AI infrastructure architecture
  • Excellent analytical, troubleshooting, communication, and technical leadership skills
  • Experience optimizing production-scale LLM inference and serving large foundation models
  • Hands-on experience with vLLM, TensorRT-LLM, DeepSpeed, Triton Inference Server, CUTLASS, FasterTransformer, or similar AI optimization frameworks
  • Knowledge of model compression, KV cache optimization, speculative decoding, and advanced inference optimization techniques
  • Experience implementing FinOps strategies for AI infrastructure cost optimization and resource management
  • Contributions to AI systems research, open-source AI infrastructure projects, patents, or technical publications
  • Familiarity with emerging AI accelerator technologies, including AMD ROCm, Intel oneAPI, or custom AI hardware

Similar Jobs

Bright Vision Technologies Logo Bright Vision Technologies

AI Systems Performance Specialist

Artificial Intelligence • Information Technology • Software • Consulting
In-Office
Huntersville, NC, USA
53 Employees
130K-180K Annually

The Aerospace Corporation Logo The Aerospace Corporation

Information Systems, IT, Data Science Intern - Summer 2027 (U.S. Person Required)

Aerospace • Artificial Intelligence • Cloud • Machine Learning • Software • Cybersecurity • Defense
Remote or Hybrid
United States
4600 Employees
23-45 Hourly

The Aerospace Corporation Logo The Aerospace Corporation

Manufacturing & Industrial Engineering Intern - Summer 2027

Aerospace • Artificial Intelligence • Cloud • Machine Learning • Software • Cybersecurity • Defense
Remote or Hybrid
United States
4600 Employees
24-45 Hourly

The Aerospace Corporation Logo The Aerospace Corporation

R&D Intern - Summer 2027 (U.S. Person Required)

Aerospace • Artificial Intelligence • Cloud • Machine Learning • Software • Cybersecurity • Defense
Remote or Hybrid
United States
4600 Employees
25-45 Hourly
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
53 Employees
Year Founded: 2020

What We Do

Bright Vision Technologies is a minority-owned organization founded in July 2020 and based in New Jersey, USA. The company specializes in delivering top-tier staffing and IT consulting services, including custom computer programming and systems design. Additionally, they are a product engineering firm with a flagship AI-powered talent intelligence and enterprise automation platform called Lumina, which helps transform IT into a strategic asset for their valued partners.

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account