Responsibilities
- Lead technical strategy for enabling and optimizing frontier-scale AI models across Microsoft AI accelerators and cloud-scale AI infrastructure.
- Architect and optimize AI systems across hardware, compilers, kernels, frameworks, runtimes, and distributed infrastructure to improve training and inference performance.
- Partner with hardware architecture teams on hardware-software co-design, influencing accelerator features, memory systems, interconnects, execution models, and future silicon roadmaps.
- Drive performance optimization across kernels, communication, memory movement, quantization, attention, mixture-of-experts (MoE), and other critical AI workloads.
- Architect distributed training and inference solutions spanning large accelerator clusters, including parallelism, communication, memory management, and scaling strategies.
- Drive model enablement and performance improvements across AI frameworks and inference technologies such as PyTorch, Triton, vLLM, SGLang, and related ecosystems.
- Lead architecture reviews and complex cross-organization technical initiatives, build alignment across engineering teams, mentor engineers, and provide technical recommendations that inform Microsoft's long-term AI infrastructure strategy.
Qualifications
Required/minimum qualifications
- Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.
- Experience partnering with hardware architecture teams on hardware-software co-design, accelerator enablement, or performance optimization.
Other Requirements
- Ability to meet Microsoft, customer, and/or government security screening requirements, including Microsoft Cloud Background Check requirements.
- 7+ years of experience developing and optimizing high-performance AI systems, kernels, or accelerator software using CUDA, ROCm, Triton, or similar programming models.
- Deep experience optimizing large-scale AI workloads, including attention, mixture-of-experts (MoE), quantization, FP8, KV-cache management, memory efficiency, or related techniques.
- Experience enabling and optimizing large language, reasoning, multimodal, or other foundation models on AI accelerators.
- Experience designing distributed training or inference systems using techniques such as tensor, pipeline, expert, or sequence parallelism.
- Deep knowledge of AI frameworks such as PyTorch and experience optimizing production-scale training or inference workloads.
- Demonstrated experience leading complex technical initiatives across multiple engineering organizations and influencing technical strategy beyond immediate team boundaries.
- Publications, patents, open-source contributions, or other recognized contributions in AI systems, distributed computing, machine learning infrastructure, or hardware acceleration.
Software Engineering IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.
Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay
This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.
Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.
Skills Required
- Bachelor's degree in Computer Science or a related technical field, or equivalent experience
- 6+ years of technical engineering experience with coding in languages including C, C++, C#, Java, JavaScript, or Python
- Experience partnering with hardware architecture teams on hardware-software co-design, accelerator enablement, or performance optimization
- Ability to meet Microsoft, customer, and/or government security screening requirements, including Microsoft Cloud Background Check requirements
- 7+ years developing and optimizing high-performance AI systems, kernels, or accelerator software using CUDA, ROCm, Triton, or similar programming models
- Experience optimizing large-scale AI workloads, including attention, mixture-of-experts, quantization, FP8, KV-cache management, or memory efficiency
- Experience enabling and optimizing large language, reasoning, multimodal, or other foundation models on AI accelerators
- Experience designing distributed training or inference systems using tensor, pipeline, expert, or sequence parallelism
- Deep knowledge of AI frameworks such as PyTorch and experience optimizing production-scale training or inference workloads
- Experience leading complex technical initiatives across multiple engineering organizations and influencing technical strategy beyond the immediate team
- Publications, patents, open-source contributions, or other recognized contributions in AI systems, distributed computing, machine learning infrastructure, or hardware acceleration
Microsoft Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Microsoft and has not been reviewed or approved by Microsoft.
-
Fair & Transparent Compensation — Pay is presented as broadly competitive overall, with clear role/level/location variation and an emphasis on using posted ranges and band information for apples-to-apples comparisons.
-
Retirement Support — Retirement benefits are described as a standout, highlighted by a strong 401(k) match structure and immediate vesting, plus additional plan features for tax-advantaged saving.
-
Parental & Family Support — Family-oriented benefits are portrayed as a meaningful strength, with substantial paid parental leave and added supports like back-up care and adoption/surrogacy assistance.
Microsoft Insights
What We Do
At Microsoft, our mission is to empower every person and every organization on the planet to achieve more. Our mission is grounded in both the world in which we live and the future we strive to create. Today, we live in a mobile-first, cloud-first world, and the transformation we are driving across our businesses is designed to enable Microsoft and our customers to thrive in this world.



.png)





