Reinforcement Learning Engineer

Posted 4 Days Ago
Be an Early Applicant
Shrewsbury, MA, USA
In-Office
100K-150K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Consulting
The Role
Design, implement, and deploy reinforcement learning systems for real and simulated environments. Develop RL algorithms, reward functions, offline and imitation learning methods, scalable distributed training infrastructure, and safety mechanisms. Train neural network policies on GPU clusters, evaluate models against adversarial and out-of-distribution cases, monitor production behavior, and collaborate with scientists and product teams to deliver impactful RL solutions.
Summary Generated by Built In
Reinforcement Learning Engineer - Remote
Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.
Job Title: Reinforcement Learning Engineer
Location: 100% Remote (U.S.)
Position Type: Full-time, Direct W2
Salary Range: $100,000–$150,000 Annually
Experience Required: 6+ years
Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Key Responsibilities
  • Design and implement reinforcement learning solutions for sequential decision-making problems in real and simulated environments.
  • Develop, calibrate, and maintain simulation environments suitable for large-scale agent training.
  • Implement and evaluate modern RL algorithms including policy gradient, actor-critic, off-policy, and offline RL methods.
  • Engineer reward functions and shaping strategies that align agent behavior with desired outcomes and safety constraints.
  • Apply offline RL and imitation learning techniques where exploration is costly or unsafe.
  • Use RLHF, DPO, and related techniques for fine-tuning large language models when relevant.
  • Build scalable training infrastructure for distributed RL, including efficient experience collection and replay systems.
  • Optimize training stability and sample efficiency through algorithmic and engineering improvements.
  • Design rigorous evaluation protocols, including out-of-distribution and adversarial test cases.
  • Implement safety mechanisms such as constraint enforcement, conservative policies, and human-in-the-loop oversight.
  • Collaborate with applied scientists and product teams to identify high-value RL use cases.
  • Monitor deployed policies and models in production for drift, regression, and unintended behaviors, building the alerting and dashboards that surface issues before they meaningfully affect users.
  • Document methodology, design decisions, and operational characteristics for internal stakeholders.
  • Stay current with RL research and translate promising techniques into production-ready solutions.

Required Qualifications
  • Master’s or PhD in Computer Science, Machine Learning, or a related field; or equivalent applied experience.
  • Six or more years of combined RL research and engineering experience.
  • Strong proficiency in Python and modern deep learning frameworks.
  • Hands-on experience with at least one major RL library or in-house RL stack.
  • Solid understanding of probability, optimization, and the theoretical foundations of RL.
  • Experience designing and tuning reward functions in non-trivial environments.
  • Familiarity with simulation environments and large-scale experience collection.
  • Experience training neural network policies on GPU clusters.
  • Strong written and verbal communication skills.
  • Track record of shipping or publishing impactful RL work.

Preferred Qualifications
  • Experience with RLHF for large language models.
  • Familiarity with multi-agent RL or hierarchical RL.
  • Exposure to robotics, control systems, or autonomous driving.
  • Publications in RL or related research venues.
  • Open-source contributions to RL libraries or environments.

How to Apply
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 505-3899. Learn more about Bright Vision Technologies at www.bvteck.com.
Bright Vision Technologies is an Equal Opportunity Employer.
 

Skills Required

  • Master’s or PhD in Computer Science, Machine Learning, or a related field, or equivalent applied experience
  • Six or more years of combined reinforcement learning research and engineering experience
  • Strong proficiency in Python and modern deep learning frameworks
  • Hands-on experience with at least one major reinforcement learning library or in-house RL stack
  • Solid understanding of probability, optimization, and reinforcement learning theory
  • Experience designing and tuning reward functions in non-trivial environments
  • Familiarity with simulation environments and large-scale experience collection
  • Experience training neural network policies on GPU clusters
  • Strong written and verbal communication skills
  • Track record of shipping or publishing impactful reinforcement learning work
  • Experience with reinforcement learning from human feedback for large language models
  • Familiarity with multi-agent or hierarchical reinforcement learning
  • Exposure to robotics, control systems, or autonomous driving
  • Publications in reinforcement learning or related research venues
  • Open-source contributions to reinforcement learning libraries or environments
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
53 Employees
Year Founded: 2020

What We Do

Bright Vision Technologies is a minority-owned organization founded in July 2020 and based in New Jersey, USA. The company specializes in delivering top-tier staffing and IT consulting services, including custom computer programming and systems design. Additionally, they are a product engineering firm with a flagship AI-powered talent intelligence and enterprise automation platform called Lumina, which helps transform IT into a strategic asset for their valued partners.

Similar Jobs

LiveKit Logo LiveKit

Research Engineer (Reinforcement Learning)

Artificial Intelligence • Cloud • Information Technology • Software
In-Office or Remote
30 Locations
83 Employees
135K-300K Annually

LiveKit Logo LiveKit

Research Engineer (Reinforcement Learning)

Artificial Intelligence • Information Technology • Internet of Things
In-Office or Remote
30 Locations
34 Employees
135K-300K Annually

Eka Robotics Logo Eka Robotics

Machine Learning / Reinforcement Learning Engineer

Artificial Intelligence • Computer Vision • Hardware • Logistics • Machine Learning • Robotics • Automation
In-Office
Boston, MA, USA
23 Employees

Eka Robotics Logo Eka Robotics

Infrastructure Engineer

Artificial Intelligence • Computer Vision • Hardware • Logistics • Machine Learning • Robotics • Automation
In-Office
Boston, MA, USA
23 Employees

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account