About the Role
- Develop high-performance GEMM, convolution, and attention kernels for AI workloads
- Design scalable JIT and codegen infrastructure for GPU kernel generation
- Implement fusion and memory-traffic optimizations to maximize hardware utilization
- Optimize mixed-precision and quantized execution paths (e.g., BF16, FP16, INT8, FP8, FP4, etc.)
- Build analytical and empirical performance models for kernel dispatch and tuning
- Profile and eliminate performance bottlenecks across oneDNN GPU primitives and runtime paths
- Co-design GPU primitives and kernel architectures for next-generation Intel GPUs
- Partner with hardware and compiler teams to shape future accelerator capabilities and software stacks
- Improve validation, benchmarking, and CI infrastructure for performance-critical GPU workloads
- Work on a global, high-impact open-source library that scales AI performance across millions of devices worldwide
- Get early access to and influence the software stack for Intel's roadmap of next-generation discrete GPUs
- Work alongside industry-leading experts in GPU compilers, hardware architecture, and performance libraries
- Enjoy a competitive package including stock programs, quarterly bonuses, robust healthcare, and highly flexible hybrid/remote working options
- A strong ownership mindset — you take initiative on complex, ambiguous technical problems and drive them to resolution
- A collaborative approach — you work effectively across hardware, compiler, and framework teams to align on shared technical goals
- A performance-driven curiosity — you are motivated by squeezing every cycle out of hardware and continuously seek deeper understanding of low-level systems
- Education: BSc, MSc, or PhD in Computer Science, Computer Engineering, Mathematics, Physics, or a highly technical related field
- Core Language: 5+ years of professional software development experience with expert-level modern C++
- Performance Optimizations: 2+ years of hands-on experience in programming and kernel optimization on GPUs (via SYCL/DPC++, OpenCL, CUDA, or HIP), or at least 5+ years of similar low-level performance optimization experience on CPUs
- Hardware Architecture: Strong foundations in computer architecture, cache hierarchies, memory subsystems, and parallel programming paradigms (e.g., multi-threading, SIMD/vectorization)
- Math Libraries: Experience developing high-performance math libraries (e.g., GEMM, convolution, reduction, or FFT kernels)
- Low-Level Tuning: Hands-on experience with GPU assembly-level tuning or compiler optimization
- Parallel APIs: Familiarity with parallel programming APIs such as OpenMP or oneTBB
- AI Workload Context: Basic understanding of deep learning primitives (e.g., forward/backward passes) to understand how library code is utilized by upstream frameworks
Job Type:Experienced HireShift:Shift 1 (United States of America)Primary Location: US, Oregon, HillsboroAdditional Locations:US, California, Santa ClaraPosting Statement:All qualified applicants will receive consideration for employment without regard to race, color, religion, religious creed, sex, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, military and veteran status, marital status, pregnancy, gender, gender expression, gender identity, sexual orientation, or any other characteristic protected by local law, regulation, or ordinance.Position of TrustN/ABenefits
We offer a total compensation package that ranks among the best in the industry. It consists of competitive pay, stock bonuses, and benefit programs which include health, retirement, and vacation. Find out more about the benefits of working at Intel.
Annual Salary Range for jobs which could be performed in the US: $195,200.00-275,580.00 USD
The range displayed on this job posting reflects the minimum and maximum target compensation for the position across all US locations. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific compensation range for your preferred location during the hiring process.
Work Model for this Role
This role will be eligible for our hybrid work model which allows employees to split their time between working on-site at their assigned Intel site and off-site. * Job posting details (such as work model, location or time type) are subject to change.*
ADDITIONAL INFORMATION: Intel is committed to Responsible Business Alliance (RBA) compliance and ethical hiring practices. We do not charge any fees during our hiring process. Candidates should never be required to pay recruitment fees, medical examination fees, or any other charges as a condition of employment. If you are asked to pay any fees during our hiring process, please report this immediately to your recruiter.Skills Required
- BSc, MSc, or PhD in Computer Science, Computer Engineering, Mathematics, Physics, or a highly technical related field
- 5+ years of professional software development experience
- Expert-level modern C++ proficiency
- 2+ years of hands-on GPU programming and kernel optimization using SYCL, DPC++, OpenCL, CUDA, or HIP
- Alternatively, 5+ years of similar low-level performance optimization experience on CPUs
- Strong foundations in computer architecture, cache hierarchies, memory subsystems, and parallel programming paradigms
- Experience developing high-performance math libraries, such as GEMM, convolution, reduction, or FFT kernels
- GPU assembly-level tuning or compiler optimization experience
- Familiarity with OpenMP or oneTBB
- Basic understanding of deep learning primitives, including forward and backward passes
Intel Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Intel and has not been reviewed or approved by Intel.
-
Leave & Time Off Breadth — Sabbaticals and paid time off are highlighted as signature elements, with an established program offering four weeks after four years or eight weeks after seven years. This distinctive time off is positioned as a meaningful part of the overall package.
-
Parental & Family Support — Paid bonding leave of 12 weeks and a New Parent Reintegration program, plus fertility benefits around $40,000 and up to $15,000 adoption reimbursement with no lifetime cap, are clearly stated. These programs are presented as standout components alongside broader family support resources.
-
Healthcare Strength — Multiple medical plan options with 2026 updates, a shift to Spring Health for EAP, and in‑network virtual medical visits covered at 100% beginning in 2026 indicate a comprehensive offering. These features signal an emphasis on robust medical access and mental health support.
Intel Insights
What We Do
Our mission is to shape the future of technology to help create a better future for the entire world, that’s the power of Intel Inside. With more ingenuity and creativity inside, our work is at the heart of countless innovations. From major breakthroughs to things that make everyday life better— they’re all powered by Intel technology. With a career at Intel, you can help make the future more wonderful for everyone.






