The Role
Optimize end-to-end AI training and inference pipelines for throughput, latency, and cost. Implement model compression, distributed training/inference optimizations, compiler-level improvements, KV cache and batching strategies, and build benchmark/regression suites. Collaborate with ML and platform teams, evaluate hardware/software offerings, and document performance playbooks.
Summary Generated by Built In
AI Performance Engineer – Remote
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title: AI Performance Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $100,000–$150,000 Annually
Experience Required: 6+ years
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Key Responsibilities
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 505-3899. Learn more about Bright Vision Technologies at www.bvteck.com.
Bright Vision Technologies is an Equal Opportunity Employer.
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title: AI Performance Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $100,000–$150,000 Annually
Experience Required: 6+ years
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Key Responsibilities
- Profile and optimize end-to-end AI training and inference pipelines for throughput, latency, and cost.
- Identify and eliminate bottlenecks across data loading, model compute, communication, and memory.
- Implement and tune quantization, sparsity, and pruning strategies to reduce model footprint and accelerate inference.
- Optimize distributed training using tensor parallelism, pipeline parallelism, FSDP, and ZeRO-style sharding.
- Tune attention implementations using FlashAttention, paged attention, and related techniques.
- Implement KV cache optimization, continuous batching, and speculative decoding for LLM serving.
- Drive compiler-level optimizations using Triton, XLA, TorchInductor, or TVM, working with the broader ML framework community to land improvements that translate into measurable end-to-end performance gains.
- Optimize data pipelines, sharding strategies, and storage access patterns for high-throughput training.
- Build and maintain rigorous benchmark suites and regression frameworks across workloads.
- Collaborate with ML and platform engineering teams to embed best practices in standard pipelines.
- Drive cost-efficiency improvements through model architecture, hardware selection, and scheduling strategies.
- Evaluate new hardware and software offerings, and advise on adoption.
- Document performance tuning playbooks and share findings broadly across engineering teams.
- Stay current with AI systems research and translate advances into production improvements.
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.
- Six or more years of experience in performance engineering, ML systems, or HPC.
- Strong proficiency in Python and C++.
- Hands-on experience optimizing deep learning workloads on modern GPUs.
- Deep understanding of distributed training and inference techniques.
- Experience with profiling tools across CPU, GPU, and distributed systems.
- Familiarity with model compression techniques and their accuracy implications.
- Strong grasp of memory hierarchies, communication primitives, and parallelism strategies.
- Excellent measurement, debugging, and analytical reasoning skills.
- Strong communication and collaboration skills.
- Experience optimizing LLM inference at production scale.
- Contributions to vLLM, TensorRT-LLM, DeepSpeed, or similar projects.
- Familiarity with custom kernel authoring in Triton or CUTLASS.
- Experience with FinOps for AI workloads.
- Publications or talks on AI systems performance.
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 505-3899. Learn more about Bright Vision Technologies at www.bvteck.com.
Bright Vision Technologies is an Equal Opportunity Employer.
Skills Required
- Bachelor's or Master's degree in Computer Science, Computer Engineering, or related field
- Six or more years of experience in performance engineering, ML systems, or HPC
- Strong proficiency in Python and C++
- Hands-on experience optimizing deep learning workloads on modern GPUs
- Deep understanding of distributed training and inference techniques
- Experience with profiling tools across CPU, GPU, and distributed systems
- Familiarity with model compression techniques and their accuracy implications
- Strong grasp of memory hierarchies, communication primitives, and parallelism strategies
- Excellent measurement, debugging, and analytical reasoning skills
- Strong communication and collaboration skills
- Experience optimizing LLM inference at production scale
- Contributions to vLLM, TensorRT-LLM, DeepSpeed, or similar projects
- Familiarity with custom kernel authoring in Triton or CUTLASS
- Experience with FinOps for AI workloads
- Publications or talks on AI systems performance
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
Bright Vision Technologies is a minority-owned organization founded in July 2020 and based in New Jersey, USA. The company specializes in delivering top-tier staffing and IT consulting services, including custom computer programming and systems design. Additionally, they are a product engineering firm with a flagship AI-powered talent intelligence and enterprise automation platform called Lumina, which helps transform IT into a strategic asset for their valued partners.







