[VMT] Platform AI Software Engineer

Posted Yesterday
Be an Early Applicant
Buenos Aires, Ciudad Autónoma de Buenos Aires, ARG
In-Office
Mid level
Software
The Role
Build and maintain the reliability layer for an AI copilot used by process engineers. Investigate agent failures, develop evaluation harnesses for prompts and model migrations, implement grounding and assumption checks, review generated outputs, and resolve performance and concurrency issues. The role requires strong Python development, production LLM experience, rigorous regression testing, feature-flagged rollouts, and autonomous debugging in a distributed team.
Summary Generated by Built In
Company Description

We are Software Mind, an awesome team of engineers who are ready to ramp up any top-notch company’s projects! Our aim? To always be one step ahead. Become part of a multicultural company in constant growth with an excellent work environment certified by Great Place To Work!

Job Description

Project - the aim you'll have

Our client builds an AI copilot for process engineers in oil refineries and chemical plants: a natural-language interface where engineers ask questions about live plant data — equipment, sensor tags, process trends — and get grounded, chart-backed answers. The users are experienced engineers who are rightly skeptical of AI: in this domain, a fabricated number or a silent wrong assumption has real cost. The product wins or loses on whether the agent can be trusted.

This role owns the trust layer of that agent inside a large, active Python codebase. It is not feature work with an LLM endpoint bolted on. The work is the mechanics of agent reliability: making the agent say "I don't know" instead of inventing, surfacing every assumption it makes so the user can correct it, grounding every claim in actual data, holding output quality through model migrations, and keeping latency acceptable while doing all of the above.

To make the day-to-day concrete, this is what the engineer currently in this seat shipped in the last four months (all of it flag-gated, in small PRs, reviewed async daily by a team spread across the US and Australia):

- An assumption auditor: detects the silent assumptions the agent makes when answering (which equipment, which time window), validates them via multi-draw consensus, and surfaces them in the UI as correctable chips — the engineer can fix an assumption and rerun the analysis.
- A grounding auditor that catches reports fabricated from empty data feeds before they reach the user.
- An adversarial reviewer sidecar that critiques generated charts for correctness before display.
- Successive frontier-model evaluations (loop behavior, directive adherence, regression on a replay harness) that decided when to flip the product's default model — including, twice, deciding NOT to flip.
- Hardening of a plant-exploration tool against hallucinating structure that the data does not support.
- A latency fix: a narration side-loop was inflating query response times; capped it and made it best-effort.
- A concurrency fix making a shared data-reset path atomic, eliminating intermittent production read errors.

If reading that list is more interesting to you than building another CRUD feature, this role is for you.

Qualifications

Expectations - the experience you need

  • Strong Python in large, shared, evolving backend codebases: you will work daily in code you didn't write, alongside people committing to it every day.
  • You have shipped an LLM-based feature to production AND built an evaluation that changed a real decision (a model choice, a prompt rollback, a killed feature).
  • Production debugging from symptom to confirmed root cause: latency spikes, concurrency errors, failures that produce no log line.
  • Prompt work treated as engineering: measured adherence, regression testing against a fixed case set — not vibes.
  • Comfort with feature-flag discipline and staged rollouts (default-off, soak, flip), small PRs, and mostly-async collaboration across US and Australia time zones.
  • High autonomy: problems arrive ambiguous ("the agent feels slow", "the engineers don't trust the numbers") and you turn them into scoped, verifiable fixes without waiting for a spec.
  • Direct, precise written English.

Nice to have

  • GCP (Vertex AI in particular); AWS/Azure acceptable.
  • Observability tooling (tracing, structured logging, latency percentiles you actually watched).
  • Experience with charting/plotting pipelines (matplotlib or similar) feeding a UI.
  • Industrial, process, or time-series data domain experience.
  • Heavy AI-tooling development workflow (Claude Code or similar) — the team works this way.

What you will do 

  • Investigate agent misbehavior reported from real customer plants and turn each case into a diagnosis, a fix, and a regression test.
  • Build and extend the evaluation harnesses that gate prompt changes and model migrations.
  • Add reliability mechanisms to the agent: assumption surfacing, grounding checks, output-quality reviewers.
  • Diagnose and resolve cross-cutting performance and concurrency issues.
  • Raise code quality in the areas you touch, within the team's review conventions.

Our Benefits 

  • Educational resources
  • Flexible schedule and Work From Anywhere
  • Referral Program
  • Supportive and chill atmosphere

We are accepting applications from LATAM countries
 

Skills Required

  • Strong Python experience in large, shared, evolving backend codebases
  • Shipped an LLM-based feature to production
  • Built an evaluation that changed a real production decision, such as a model choice, prompt rollback, or feature cancellation
  • Production debugging experience identifying confirmed root causes for latency, concurrency, and silent failures
  • Prompt engineering experience using measured adherence and regression testing
  • Experience with feature flags and staged rollouts
  • Ability to work autonomously on ambiguous problems
  • Direct, precise written English
  • GCP experience, particularly Vertex AI
  • Observability tooling experience, including tracing, structured logging, and latency monitoring
  • Charting or plotting pipeline experience, such as Matplotlib
  • Industrial, process, or time-series data domain experience
  • Experience with AI development tools such as Claude Code

Software Mind Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Software Mind and has not been reviewed or approved by Software Mind.

  • Fair & Transparent Compensation Pay is considered competitive for core hiring markets, with “good salary” cited in multiple locales. Public salary snapshots provide a baseline that helps candidates assess offers and negotiations.
  • Flexible Benefits Remote or hybrid options are prominently highlighted, and a remote‑work program is publicly noted alongside positively cited work‑from‑home experiences. Flexibility around schedules and location is presented as part of the package.
  • Wellbeing & Lifestyle Benefits Private medical care, language classes, sports/fitness support, and learning initiatives are listed for several Central/Eastern European locations, with occasional workation perks promoted. These lifestyle‑oriented offerings complement base pay and can enhance perceived total rewards.

Software Mind Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Kraków
1,000 Employees
Year Founded: 1999

What We Do

Software Mind is a global digital transformation partner with operations throughout Europe, the US and LATAM. Driven by tech and empowered by people, we provide companies with software engineers and autonomous, cross-functional development teams who manage software life cycles from ideation to release and beyond. For over 20 years we’ve been enriching organizations with the talent they need to boost scalability, drive dynamic growth and bring disruptive ideas to life. Our top-notch engineering teams combine ownership with leading technologies, including cloud, AI, data science and embedded software to accelerate digital transformations and boost software delivery. A culture, driven by trust, that embraces openness, craves more and acts with respect enables our experts to create evolutive solutions that support scale-ups, unicorns and enterprise-level companies around the world.

Similar Jobs

Mastercard Logo Mastercard

Manager, Account Management

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Remote or Hybrid
Buenos Aires, Ciudad Autónoma de Buenos Aires, ARG
38800 Employees

ZS Logo ZS

Senior Project Manager

Artificial Intelligence • Healthtech • Professional Services • Analytics • Consulting
Hybrid
Ciudad Autónoma de Buenos Aires, ARG
15000 Employees
6-10 Annually

Pfizer Logo Pfizer

Pasante Asuntos Regulatorios | Argentina - Buenos Aires

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office
Buenos Aires, Ciudad Autónoma de Buenos Aires, ARG
121990 Employees

Pfizer Logo Pfizer

Analista de Asuntos Regulatorios - Plazo fijo 1 año

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office
Buenos Aires, Ciudad Autónoma de Buenos Aires, ARG
121990 Employees

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account