Software Engineer -Gen AI Inferencing | Onsite - (Dallas, Charlotte, New York)

Posted 2 Days Ago
2 Locations
In-Office or Remote
44K-154K Annually
Senior level
Agency • Information Technology
The Role
Designs, builds, tests, deploys, and operates reusable GenAI inference and RAG toolkits. Develops scalable Python-based systems, model-serving frameworks, APIs, automation, CI/CD pipelines, and performance-optimized production deployments. Collaborates with product, data, and engineering teams; contributes to architecture, requirements, testing, compliance, and agile delivery. Mentors engineers, automates releases, prototypes emerging technologies, and supports GenAI platform capabilities including inferencing, observability, policy management, and MCP use cases.
Summary Generated by Built In

"Required qualifications:
5+ years OOP in Python/Scala/Java programming experience with expert level development skills
Experience with AI/ML/GenAI Lifecycle Management and Development and its Ecosystem. 
-Hands on experience building frameworks using MLOps, Fine – Tuning techniques, Inference Frameworks
-Experience with deploying models using vLLM/Triton Inference Server in containers in production with automation. 
-Performs Continuous Integration and Continuous Development (CI-CD) activities.
Performance Tuning those models and deployment to provide higher throughput.
-Track record of maintaining large scale Python/Unix based systems.
-Hands on experience and knowledge generative AI RAG process for various use cases, including chunking, embedding, retrieval, reranking and summarization.
-Hands-on experience in application development in one or more areas MongoDB, Redis, Angular/React Frameworks, Containerization, Building API based application leveraging FAST API services, JWT Integration, API Gateway
Develop efficient utilities, automation frameworks, data science platforms that can be utilized across multiple Data Science teams for AI/ML and GenAI work.
Working in large sized teams that collaboratively develop on a shared multi-repo codebase using IDEs (e.g. VS Code rather than Jupyter Notebooks), Continuous Integration (CI), Continuous Deployment (CD) and Continuous Testing
Strong automation, scripting, and Python development skills. Hands-on DevOps experience with one or more of the following enterprise development tools: Version Control (GIT/Bitbucket), Build Orchestration (Jenkins), Code Quality (SonarQube and pytest Unit Testing), Artifact Management (Artifactory) and Deployment (Ansible)"
 

"Position Summary
- This position is focused on design, build, and operate of reusable toolkits for Gen AI RAG capabilities.
-This job is responsible for developing and delivering complex requirements to accomplish business goals. Key responsibilities of the job include ensuring that software is developed to meet functional, non-functional and compliance requirements, and solutions are well designed with maintainability/ease of integration and testing built-in from the outset. Job expectations include a strong knowledge of development and testing practices common to the industry and design and architectural patterns."
 

"Responsibilities:
Codes solutions and unit test to deliver a requirement/story per the defined acceptance criteria and compliance requirements
Designs, develops, and modifies architecture components, application interfaces, and solution enablers while ensuring principal architecture integrity is maintained
Mentors other software engineers and coach team on Continuous Integration and Continuous Development (CI-CD) practices and automating tool stack
Executes story refinement, definition of requirements, and estimating work necessary to realize a story through the delivery lifecycle
Performs spike/proof of concept as necessary to mitigate risk or implement new ideas
Automates manual release activities
Designs, develops, and maintains automated test suites (integration, regression, performance)
Utilizes multiple architectural components (across data, application, business) in design and development of client requirements
Manage multiple priorities, and simultaneously engage with multiple teams.
Participates in estimating work necessary to realize a story/requirement through the delivery lifecycle.
Be vocal and actively participate in all session with business stakeholders and agile teams.
Collaborate with product teams, data analysts and data scientists to design and build solutions."

"Desired Qualifications
Experience building & deploying Gen AI inferencing platform with open-source toolsets, building inferencing & servicing capabilities (AI Gateway, Policy store, Observability) for RAG/ MCP use cases etc.
Hands on experience on driving and maintaining a culture of quality, innovation, and experimentation.
Research on new tools and capabilities for better UI and UX for advanced analytics platform, quick prototype and demonstrate the features and capabilities, and participate on various user forums."


Compensation, Benefits and Duration

Minimum Compensation: USD 44,000
Maximum Compensation: USD 154,000
Compensation is based on actual experience and qualifications of the candidate. The above is a reasonable and a good faith estimate for the role.
Medical, vision, and dental benefits, 401k retirement plan, variable pay/incentives, paid time off, and paid holidays are available for full-time employees.
This position is not available for independent contractors
No applications will be considered if received more than 120 days after the date of this post

Skills Required

  • 5+ years of object-oriented programming experience with Python, Scala, or Java
  • Expert-level development skills
  • Experience with AI/ML/GenAI lifecycle management, development, and ecosystem
  • Hands-on experience building MLOps, fine-tuning, and inference frameworks
  • Experience deploying models with vLLM or Triton Inference Server in production containers with automation
  • Experience with CI/CD activities
  • Experience performance tuning models and deployments for higher throughput
  • Track record maintaining large-scale Python/Unix-based systems
  • Hands-on knowledge of generative AI RAG processes, including chunking, embedding, retrieval, reranking, and summarization
  • Application development experience with MongoDB, Redis, Angular or React, containerization, FastAPI, JWT, or API Gateway
  • Strong automation, scripting, and Python development skills
  • Hands-on DevOps experience with enterprise development tools such as Git/Bitbucket, Jenkins, SonarQube, pytest, Artifactory, or Ansible
  • Experience working on shared multi-repository codebases with continuous integration, deployment, and testing
  • Experience building and deploying GenAI inferencing platforms with open-source toolsets
  • Experience building inferencing and servicing capabilities such as AI Gateway, policy store, and observability for RAG/MCP use cases
  • Experience driving a culture of quality, innovation, and experimentation
  • Experience researching new tools, prototyping UI/UX capabilities, and demonstrating advanced analytics platform features
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: London
5,017 Employees
Year Founded: 2007

What We Do

Photon.com has emerged as one of the world’s largest and fastest-growing Digital Agencies. We work with 40% of the Fortune 100 on their Digital initiatives and are known for our ability to integrate Strategy Consulting, Creative Design, and Technology at scale. Please visit www.photon.com to learn more about us, how we work, and our customer case studies. Digital Transformation Starts Here.

Similar Jobs

Remote
USA
123 Employees

Spectrum Logo Spectrum

Director, Enterprise New Business Sales (RapidScale) (Remote)

Information Technology • Internet of Things • Mobile • On-Demand • Software
In-Office or Remote
Raleigh, NC, USA
100000 Employees

Capital One Logo Capital One

Referral & Lifecycle Marketing Lead, Capital One Shopping (Remote-Eligible)

Fintech • Machine Learning • Payments • Software • Financial Services
Remote or Hybrid
3 Locations
55000 Employees
210K-287K Annually

Capital One Logo Capital One

Lead Software Engineer

Fintech • Machine Learning • Payments • Software • Financial Services
Remote or Hybrid
McLean, VA, USA
55000 Employees
179K-225K Annually

Similar Companies Hiring

Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account