Rational Dynamics builds customized AI reasoning systems for tasks of high cognitive complexity.
Our initial market is the world’s leading institutional asset owners. We work very closely with these customers to create specialized, rigorous benchmark datasets encompassing their most valuable and difficult knowledge work. Then we use the benchmarks to construct agentic large reasoning models, applying the same rigor to prove that the models correctly do the work. Customers access the models through a tailored application service, making their most skilled, expensive workers dramatically more productive.
We are an early-stage startup. Our founders previously started Voleon, now one of the world’s largest systematic investment managers, and recognized as a longstanding industry leader in applied machine learning. They bring to Rational Dynamics the same research discipline and data-driven focus that succeeded in the unforgiving, high-stakes setting of financial markets.
About the roleInstitutional investors want AI for their hardest analytical work, running inside their cloud accounts and regions, under their security and compliance rules. Rational Dynamics builds that AI. Infrastructure decides whether we can ship it: each client brings a different cloud, network, identity provider, and security review. You'll be working across client deployments, internal platform, and the research infrastructure, collaborating with product development and science teams.
What you'll doDeploy our cloud-native platform into client clouds on AWS, GCP, and Azure, including multi-region and data-residency setups. Answer the security and architecture questions client IT teams raise
Co-own existing infrastructure in AWS and GCP: EKS, Aurora Postgres, Cloud Run, networking, IAM, secrets, and GitOps delivery
Run the research platform for evals and training: proprietary, open-source, managed services such as Vertex AI, SageMaker, and Bedrock, with GPU compute and storage infrastructure
Build CI/CD, observability, and alerting that let a small team ship daily and catch problems before clients do
Stand up rapid prototypes and harden existing prototypes for production. Give Forward Deployed Engineers tooling to do the same in client environments
Share compliance (e.g. SOC 2, ISO 27001, CAIQ) and security work so audit evidence comes from how the platform runs
Contribute to platform development, MLOps, and internal GenAI tools
Manage infrastructure as code in Terraform, write down the trade-offs behind each decision, and track reliability, uptime, and cost
7+ years designing and running production cloud infrastructure, 5+ of them at Senior level or above
Depth in AWS or GCP, and hands-on experience with both
Strong Kubernetes, Terraform, and CI/CD, including GitOps with ArgoCD or similar
Cloud networking and security: VPCs, private connectivity, IAM, zero-trust access, encryption, secrets, and multi-region deployments with data-residency requirements
Deployments into enterprise customer clouds in finance, healthcare, defense, or another regulated industry
Work under SOC 2, ISO 27001, PCI DSS, or a similar control framework
Core infrastructure built from scratch at an early-stage company, or owned end to end on a small team
Python and Bash for automation, and daily use of agentic coding tools such as Claude Code
Clear communication about designs and trade-offs with engineers, scientists, and client security teams, and comfort with shifting priorities
Azure (AKS, VNets, Entra ID, ML/AI) or Oracle Cloud (OCI)
ML and GPU infrastructure: Kubeflow, Vertex AI, SageMaker, Bedrock, MLflow, Ray, KServe, Slurm
Infrastructure for LLM and agent systems: model serving, vector databases, and training and eval pipelines
Workflow tools such as Temporal, Airflow, Argo Workflows, Tekton
Full-stack or agentic application development, or ML and RL research
Go or another systems language
A record of running production systems against reliability and uptime targets
If you have a great candidate in mind for this role and would like to have the potential to earn $7,500 to $15,000 if your referred candidate is successfully hired and employed by Rational Dynamics, please use this form to submit your referral. For more details regarding eligibility, terms and conditions please make sure to review the Rational Dynamics Referral Bonus Program.
Equal Opportunity EmployerRational Dynamics is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.
Skills Required
- 7+ years designing and running production cloud infrastructure
- 5+ years at Senior level or above
- Deep expertise in AWS or GCP with hands-on experience with both
- Strong Kubernetes, Terraform, and CI/CD experience, including GitOps with ArgoCD or similar
- Experience with cloud networking and security, including VPCs, private connectivity, IAM, zero-trust access, encryption, secrets, and multi-region deployments
- Experience deploying into enterprise customer clouds in finance, healthcare, defense, or another regulated industry
- Experience working under SOC 2, ISO 27001, PCI DSS, or similar control frameworks
- Experience building core infrastructure from scratch at an early-stage company or owning infrastructure end to end on a small team
- Python and Bash automation experience
- Daily use of agentic coding tools such as Claude Code
- Clear communication with engineers, scientists, and client security teams; comfort with shifting priorities
- Experience with Azure, including AKS, VNets, Entra ID, or ML/AI
- Experience with Oracle Cloud OCI
- Experience with ML and GPU infrastructure, including Kubeflow, Vertex AI, SageMaker, Bedrock, MLflow, Ray, KServe, or Slurm
- Experience with infrastructure for LLM and agent systems, including model serving, vector databases, and training or evaluation pipelines
- Experience with Temporal, Airflow, Argo Workflows, or Tekton
- Full-stack, agentic application development, ML, or RL research experience
- Go or another systems programming language
- Record of running production systems against reliability and uptime targets
What We Do
Our initial market is the world’s leading institutional asset owners. We work very closely with these customers to create specialized, rigorous benchmark datasets encompassing their most valuable and difficult knowledge work. Then we use the benchmarks to construct agentic large reasoning models, applying the same rigor to prove that the models correctly do the work. Customers access the models through a tailored application service, making their most skilled, expensive workers dramatically more productive.
.png)







