AI/ML Software Engineer-Task Creator (RL Environments) (Contract)

Reposted 11 Days Ago
Hiring Remotely in United States
Remote
Entry level
Artificial Intelligence • HR Tech • Software • Generative AI
The Role
Create end-to-end computer-use tasks for AI agents in Linux desktop environments. Design realistic multi-application scenarios, prepare files and instructions, write Python setup and automated scoring scripts, test tasks, and document them. The contract requires daily check-ins, rapid feedback, and availability throughout a three- to four-week engagement. Experience with AI tools, AI/ML projects, Linux VMs, Git, reinforcement learning environments, or computer-use benchmarks is relevant.
Summary Generated by Built In

Careerflow Human Data Labs partners with AI companies to bring real-world professional expertise into their products.

About the work

We build hard, realistic tasks used to test how well AI agents do real work on a computer. Each task is a small scenario running on a Linux desktop — real files, real applications, real documents — plus a program that automatically checks whatever the agent produced.

You create those tasks end to end. The bar: hard for a top AI agent, straightforward for a competent human.

Responsibilities

· Design a realistic multi-step scenario across a few applications.

· Put together the files it needs — spreadsheets, documents, emails, data. Sometimes supplied to you, sometimes built by you.

· Write the instruction the AI agent receives: clear, complete, no giveaways.

· Write Python to set up the environment and to score the result automatically.

· Test it, run it, and hand over short documentation.

What we are looking for

· Python — comfortable writing and debugging real scripts.

· Heavy AI user — you work with AI coding tools and models every day and get genuinely good output from them.

· Some AI / ML background — worked on AI projects, at an AI company, or with agents and evaluations.

· Ubuntu and virtual machines — basic comfort. Everything runs on hosted Linux VMs, so you should be fine on a terminal and inside a Linux desktop.

· Careful and detailed — an unclear instruction or a sloppy check breaks the task.

· Git basics.

A strong plus

· You have built RL environments before — anything where an agent acts and a program scores the outcome.

· You have built tasks or benchmarks for computer-use agents (agents that control a real desktop, browser, or operating system).

· You know OSWorld or similar computer-use benchmarks. Not required, but it means you will be productive on day one.

Payment and availability

· $50 USD per task, paid once the task is accepted and approved in review.

· One free revision round if a task needs fixes.

· No cap on how many tasks you deliver — throughput is up to you.

· Must be available to start the next day and stay through the 3–4 week window.

· Flexible hours, fully async. We ask for a daily check-in and feedback turnaround within about a day.

Process

Short application → the screening questions below → a quick call → one paid pilot task ($50 on acceptance) → onboard. Decisions within 48 hours.

Engagement

Short-term contract, 3 to 4 weeks

Start

Immediately — the day after selection

Location

Fully remote, India Only

Hours

Flexible — we measure delivered tasks, not hours

Pay

$50 USD per accepted and approved task

Skills Required

  • Comfortable writing and debugging Python scripts
  • Daily use of AI coding tools and AI models with strong ability to obtain useful output
  • Some AI or machine learning project experience, or experience with agents and evaluations
  • Basic comfort with Ubuntu, Linux terminals, and virtual machines
  • Careful and detail-oriented approach to writing instructions and automated checks
  • Basic Git knowledge
  • Availability to start the day after selection
  • Availability throughout the three- to four-week contract window
  • Prior experience building reinforcement learning environments
  • Experience creating tasks or benchmarks for computer-use agents
  • Knowledge of OSWorld or similar computer-use benchmarks
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company

What We Do

Careerflow.ai is an AI-powered career management platform and 'career copilot' dedicated to helping job seekers land their dream jobs. The company provides a comprehensive end-to-end toolkit featuring an AI resume builder, LinkedIn profile optimizer, and job tracking tools. By streamlining the application process and optimizing professional profiles, Careerflow helps users navigate the competitive job market and get hired at top tech and startup companies faster.

Similar Jobs

Dynatrace Logo Dynatrace

Analyst Relations Manager

Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Big Data Analytics • Automation
Remote or Hybrid
Boston, MA, USA
5600 Employees
96K-120K Annually

Zeta Global Logo Zeta Global

Paid Search Manager

AdTech • Artificial Intelligence • Marketing Tech • Software • Analytics
Easy Apply
Remote or Hybrid
United States
2429 Employees
80K-90K Annually

JPMorganChase Logo JPMorganChase

Controller

Financial Services
Remote or Hybrid
2 Locations
289097 Employees
Remote or Hybrid
3 Locations
289097 Employees

Similar Companies Hiring

Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account