Reinforcement Learning Engineer

Sorry, this job was removed at 02:23 p.m. (UTC) on Thursday, Aug 13, 2026
Be an Early Applicant
Lexington, MA, USA
In-Office
100K-150K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Consulting
The Role
Design, implement, and productionize reinforcement learning solutions and simulations. Develop and evaluate modern RL algorithms (policy gradient, actor-critic, off-policy, offline RL), engineer reward functions, build scalable distributed training infrastructure, apply RLHF/DPO for LLM fine-tuning, ensure safety and monitoring in production, and collaborate with product teams to translate research into deployed systems.
Summary Generated by Built In
Reinforcement Learning Engineer - Remote
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title: Reinforcement Learning Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $100,000–$150,000 Annually
Experience Required: 6+ years
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Key Responsibilities
  • Design and implement reinforcement learning solutions for sequential decision-making problems in real and simulated environments.
  • Develop, calibrate, and maintain simulation environments suitable for large-scale agent training.
  • Implement and evaluate modern RL algorithms including policy gradient, actor-critic, off-policy, and offline RL methods.
  • Engineer reward functions and shaping strategies that align agent behavior with desired outcomes and safety constraints.
  • Apply offline RL and imitation learning techniques where exploration is costly or unsafe.
  • Use RLHF, DPO, and related techniques for fine-tuning large language models when relevant.
  • Build scalable training infrastructure for distributed RL, including efficient experience collection and replay systems.
  • Optimize training stability and sample efficiency through algorithmic and engineering improvements.
  • Design rigorous evaluation protocols, including out-of-distribution and adversarial test cases.
  • Implement safety mechanisms such as constraint enforcement, conservative policies, and human-in-the-loop oversight.
  • Collaborate with applied scientists and product teams to identify high-value RL use cases.
  • Monitor deployed policies and models in production for drift, regression, and unintended behaviors, building the alerting and dashboards that surface issues before they meaningfully affect users.
  • Document methodology, design decisions, and operational characteristics for internal stakeholders.
  • Stay current with RL research and translate promising techniques into production-ready solutions.

Required Qualifications
  • Master’s or PhD in Computer Science, Machine Learning, or a related field; or equivalent applied experience.
  • Six or more years of combined RL research and engineering experience.
  • Strong proficiency in Python and modern deep learning frameworks.
  • Hands-on experience with at least one major RL library or in-house RL stack.
  • Solid understanding of probability, optimization, and the theoretical foundations of RL.
  • Experience designing and tuning reward functions in non-trivial environments.
  • Familiarity with simulation environments and large-scale experience collection.
  • Experience training neural network policies on GPU clusters.
  • Strong written and verbal communication skills.
  • Track record of shipping or publishing impactful RL work.

Preferred Qualifications
  • Experience with RLHF for large language models.
  • Familiarity with multi-agent RL or hierarchical RL.
  • Exposure to robotics, control systems, or autonomous driving.
  • Publications in RL or related research venues.
  • Open-source contributions to RL libraries or environments.

How to Apply
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 505-3899. Learn more about Bright Vision Technologies at www.bvteck.com.
Bright Vision Technologies is an Equal Opportunity Employer.
 

Skills Required

  • Master's or PhD in Computer Science, Machine Learning, or related field; or equivalent applied experience.
  • Six or more years of combined RL research and engineering experience.
  • Strong proficiency in Python and modern deep learning frameworks.
  • Hands-on experience with at least one major RL library or in-house RL stack.
  • Solid understanding of probability, optimization, and theoretical foundations of RL.
  • Experience designing and tuning reward functions in non-trivial environments.
  • Familiarity with simulation environments and large-scale experience collection.
  • Experience training neural network policies on GPU clusters.
  • Strong written and verbal communication skills.
  • Track record of shipping or publishing impactful RL work.
  • Experience with RLHF for large language models.
  • Familiarity with multi-agent RL or hierarchical RL.
  • Exposure to robotics, control systems, or autonomous driving.
  • Publications in RL or related research venues.
  • Open-source contributions to RL libraries or environments.

Similar Jobs

Bright Vision Technologies Logo Bright Vision Technologies

Reinforcement Learning Engineer

Artificial Intelligence • Information Technology • Software • Consulting
In-Office
2 Locations
53 Employees
100K-150K Annually

Eka Robotics Logo Eka Robotics

Machine Learning / Reinforcement Learning Engineer

Artificial Intelligence • Computer Vision • Hardware • Logistics • Machine Learning • Robotics • Automation
In-Office
Boston, MA, USA
23 Employees

Eka Robotics Logo Eka Robotics

Infrastructure Engineer

Artificial Intelligence • Computer Vision • Hardware • Logistics • Machine Learning • Robotics • Automation
In-Office
Boston, MA, USA
23 Employees

Percepta AI Logo Percepta AI

Scientist

Artificial Intelligence • Software
In-Office
2 Locations
3 Employees
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
53 Employees
Year Founded: 2020

What We Do

Bright Vision Technologies is a minority-owned organization founded in July 2020 and based in New Jersey, USA. The company specializes in delivering top-tier staffing and IT consulting services, including custom computer programming and systems design. Additionally, they are a product engineering firm with a flagship AI-powered talent intelligence and enterprise automation platform called Lumina, which helps transform IT into a strategic asset for their valued partners.

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account