Staff PlatformOps Engineer (Emerging AI)

Posted 2 Days Ago
Be an Early Applicant
Seattle, WA, USA
Hybrid
198K-246K Annually
Expert/Leader
eCommerce • Fintech • Logistics • Software • Transportation • Big Data Analytics
North America's Largest On-demand Freight Marketplace
The Role
Own the architecture and operation of DAT’s shared AI and machine learning platform. Build model gateways, inference services, feature and vector storage, evaluation systems, observability, governance, and cost controls on AWS. Lead distributed-systems and cloud architecture, production incident response, engineering standards, cross-functional alignment, and mentorship. Support model training, serving, monitoring, safety, and continuous improvement across AI teams.
Summary Generated by Built In

About DAT

DAT Freight & Analytics is an award-winning employer of choice and a next-generation SaaS technology company that has been at the leading edge of freight and logistics innovation for nearly five decades. Founded in 1978, DAT operates the largest freight marketplace in North America — processing 250 million+ load posts annually and maintaining one of the largest repositories of freight market transaction data in the world. On a defined path to $1 billion in revenue, DAT deploys a suite of software solutions, machine learning models, and intelligent automation tools that help brokers, carriers, and shippers price freight accurately, source capacity, reduce risk, and operate more efficiently. With nearly 700 teammates across offices in Denver, CO; Portland, OR; Seattle, WA; Springfield, MO; Toronto, ON; and Bangalore, India, DAT combines the credibility of a multi-decade market leader with the drive of a company that is not done disrupting the industry it helped build. For more information, visit www.DAT.com

 

Job Application Deadline:  10/31/2026


The Opportunity

As a Staff AI Platform Engineer, you'll build and own the platform that every AI and machine learning workload at DAT runs on. Freight is an uncertain business, and the models we ship reduce that uncertainty: rate forecasts, load-to-truck matching, document extraction, fraud signals, and the agentic workflows our brokers and carriers use to move freight faster. None of that reaches a customer without a platform that makes training, serving, evaluating, and monitoring models routine instead of heroic.

You'll set the architectural direction for that platform, from the model gateway and inference layer through feature and vector storage, evaluation harnesses, and production observability. This is a highly visible role for an engineer who wants their work multiplied across every AI team in the company.

What You'll Do

  • Platform Ownership: Design, build, and operate the shared services that engineers and business users use to ship models: a model gateway for LLM access, inference endpoints for real-time and batch scoring, feature storage, vector search, and a common SDK.
  • Technical Leadership: Lead architecture for large-scale AI systems, write the design documents, and drive alignment across Product, and Engineering and Business users  on how models get built and shipped at DAT.
  • Cloud Architecture: Architect and run scalable, reliable AI infrastructure on AWS, including Bedrock, SageMaker, EKS, Redpanda, MSK (Kafka), Lambda, and S3, all defined in Pulumi..
  • Evaluation and Quality: Build the offline and online evaluation systems that tell us whether a model or prompt change is an improvement, including regression suites, LLM-as-judge pipelines, A/B and shadow testing, and drift detection.
  • Cost and Performance: Own inference cost and latency as first-class metrics. Right-size GPU and serverless capacity, tune batching, caching, and quantization, and give teams clear visibility into what their workloads cost.
  • Safety and Governance: Implement guardrails, prompt and output logging, PII handling, access controls, and model and dataset lineage so AI systems meet our security and customer data commitments.
  • Best Practices: Drive the adoption of modern software development practices across AI work, including automated testing, code reviews, CI/CD pipelines, and infrastructure-as-code.
  • Mentorship: Mentor engineers and data scientists on production ML and distributed systems, and raise the operational bar of every team that builds on the platform.
  • Incident Management: Lead the response and resolution for complex production incidents involving AI services, perform root cause analysis, and implement preventative measures.

The Skills and Experience You'll Bring

  • 10+ years of experience in software engineering, including significant time as a senior or staff engineer owning production infrastructure that other engineering teams depend on.
  • 4+ years building and operating machine learning or AI platforms, with models you've taken from notebook to production traffic and then kept healthy.
  • Strong development language skills ideally using  TypeScript/Node.js, Java, or Go, with the software engineering discipline to build libraries other teams adopt willingly.
  • Extensive AWS experience with production systems, ideally including EKS/Kubernetes, Lambda, Redpanda, MSK/Kafka, S3, and Secrets Manager, plus infrastructure-as-code with Pulumi, Terraform or similar.
  • Hands-on experience serving models at scale: containerized inference, GPU scheduling, autoscaling, batching, and the tradeoffs between real-time, streaming, and batch scoring.
  • Practical experience with LLM systems in production, including retrieval-augmented generation, vector databases, prompt and context management, tool-calling or agent frameworks, and managed model APIs such as Amazon Bedrock.
  • Proven experience building evaluation and monitoring for models, not just services: offline eval sets, online metrics, drift and data quality checks, and a clear definition of what "regression" means for a model.
  • Strong background in event-driven and distributed systems using Redpanda, AWS MSK or Kafka
  • Fluency with observability and on-call health, designing dashboards, APM traces, logs, and alerts, defining SLOs, and using these tools to drive down incident frequency and MTTR.
  • Experience establishing and raising engineering standards for a team: PR and testing guidelines, deployment validation checklists, model release criteria, and post-incident review practices.
  • Demonstrated technical leadership: leading cross-team designs in ambiguous problem spaces, writing clear system design documents and operations runbooks, and driving alignment across Engineering, Product, and Operations.
  • Proven mentoring track record, especially helping engineers and data scientists grow in system design, observability, and operational excellence.
  • Excellent communication skills, with the ability to explain trade-offs and system behavior clearly to engineers, product managers, and non-technical stakeholders.

We'd be Extra Excited if You Have

  • Experience with model fine-tuning, distillation, or parameter-efficient training, and a clear point of view on when it beats prompting.
  • Background in marketplace, pricing, forecasting, or document extraction problems.
  • Freight, logistics, or supply chain domain experience.


Why DAT?
DAT is an award winning employer of choice.

For starters, we have a hybrid work environment, but we also know what makes a great workplace. We have a time-tested and resolute set of operating values predicated on integrity, mutual respect, open communication, and executing with excellence. These values inform our strategic vision as much as any one of our products does. We’ve been an employer of choice in the Portland metropolitan area for four decades, and within one year of opening our Denver office, DAT was #26 on Built In Colorado’s 100 Best Places to Work In Colorado.

  • Medical, Dental, Vision, Life, and AD&D insurance
  • Parental Leave
  • Flexible Vacation Time (FVT)
  • An additional 10 holidays of paid time off per calendar year
  • 401k matching (immediately vested)
  • Employee Stock Purchase Plan
  • Short- and Long-term disability sick leave
  • Flexible Spending Accounts
  • Health Savings Accounts
  • Employee Assistance Program
  • Additional programs - Employee Referral, Internal Recognition, and Wellness
  • Free TriMet transit pass (Beaverton Office)
  • Competitive salary and benefits package
  • Work on impactful projects in a cutting-edge environment
  • Collaborative and supportive team culture
  • Opportunity to make a real difference in the trucking industry
  • Employee Resource Groups


For Washington-based candidates, in compliance with the Washington State Pay Transparency Law, the salary range for this role is $198,000.00 - $246,000.00 + target bonus.  DAT considers factors such as scope and responsibilities of the position, candidate's work experience, education and training, core skills, internal equity, and market and business elements when extending an offer.

 

DAT embraces the value of a diverse workforce, and believes it is a core strength of our company that we encourage those values in every DAT employee, at every level of our organization, regardless of tenure or rank. We provide equal employment opportunities (EEO) to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, disability, genetic information, marital status, amnesty, or status as a covered veteran in accordance with applicable federal, state, and local laws.

Equal Opportunity Employer/Protected Veterans/Individuals with Disabilities

The contractor will not discharge or in any other manner discriminate against employees or applicants because they have inquired about, discussed, or disclosed their own pay or the pay of another employee or applicant. However, employees who have access to the compensation information of other employees or applicants as a part of their essential job functions cannot disclose the pay of other employees or applicants to individuals who do not otherwise have access to compensation information, unless the disclosure is (a) in response to a formal complaint or charge, (b) in furtherance of an investigation, proceeding, hearing, or action, including an investigation conducted by the employer, or (c) consistent with the contractor’s legal duty to furnish information. 41 CFR 60-1.35(c)

#LI-RF1

#LI-hybrid

Skills Required

  • 10+ years of software engineering experience, including significant senior or staff-level ownership of production infrastructure
  • 4+ years building and operating machine learning or AI platforms
  • Experience taking machine learning models from notebooks to production traffic and maintaining them
  • Strong software development skills in TypeScript/Node.js, Java, or Go
  • Extensive AWS production experience
  • Experience with EKS/Kubernetes, Lambda, Redpanda, MSK/Kafka, S3, and Secrets Manager
  • Experience with infrastructure-as-code using Pulumi, Terraform, or similar
  • Experience serving models at scale, including containerized inference, GPU scheduling, autoscaling, batching, and real-time, streaming, and batch scoring
  • Production experience with LLM systems, retrieval-augmented generation, vector databases, prompt and context management, tool-calling or agent frameworks, and managed model APIs
  • Experience building model evaluation and monitoring systems, including offline evaluations, online metrics, drift detection, data quality checks, and regression definitions
  • Strong background in event-driven and distributed systems using Redpanda, AWS MSK, or Kafka
  • Experience with observability, on-call operations, dashboards, APM traces, logs, alerts, SLOs, and reducing MTTR
  • Experience establishing engineering standards, testing guidelines, deployment validation, model release criteria, and post-incident reviews
  • Demonstrated technical leadership in cross-team system design, documentation, runbooks, and alignment
  • Proven mentoring experience with engineers and data scientists
  • Excellent communication skills with technical and non-technical stakeholders
  • Experience with model fine-tuning, distillation, or parameter-efficient training
  • Experience in marketplace, pricing, forecasting, or document extraction problems
  • Freight, logistics, or supply chain domain experience

DAT Freight & Analytics Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about DAT Freight & Analytics and has not been reviewed or approved by DAT Freight & Analytics.

  • Leave & Time Off Breadth Time off appears comparatively strong, with generous PTO and paid holidays highlighted as a standout part of the overall package. Flexible schedules and hybrid/remote options further reinforce the sense of meaningful time-off and flexibility support.
  • Retirement Support Retirement support is positioned as a strength via 401(k) matching and immediate vesting on the match. An employee stock purchase plan is also presented as an additional long-term wealth-building lever.
  • Healthcare Strength Healthcare coverage is described as broad, spanning medical, dental, vision, disability, and mental health support. The depth of plan options is framed as a meaningful benefit even when overall compensation sentiment is less favorable.

DAT Freight & Analytics Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Denver, CO
700 Employees
Year Founded: 1978

What We Do

DAT Freight & Analytics operates DAT One, North America’s largest truckload freight marketplace; Convoy Platform, an automated freight-matching technology; DAT iQ, the industry’s leading freight data analytics service; Trucker Tools, the leader in load visibility; and DAT Outgo, the freight financial services platform. Shippers, transportation brokers, carriers, news organizations, and industry analysts rely on DAT for market trends and data insights, informed by nearly 700,000 daily load posts and a database exceeding $1 trillion in freight market transactions. Founded in 1978, DAT is a business unit of Roper Technologies (Nasdaq: ROP), a constituent of the Nasdaq 100, S&P 500, and Fortune 1000. Headquartered in Beaverton, Oregon, with offices in Seattle, Denver, Springfield & Bangalore, DAT continues to set the standard for innovation in the trucking and logistics industry.

Why Work With Us

We pioneered freight technology and haven't stopped disrupting since. Our SaaS platform powers the supply chain that moves goods across America every day. What sets us apart: a teammate-first culture where we thrive on challenges and driving impact every day.

Gallery

Gallery

Similar Jobs

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Inside Sales Representative

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
12 Locations
40000 Employees
45K-100K Annually
Hybrid
Seattle, WA, USA
205000 Employees
23-31 Hourly

MetLife Logo MetLife

Sr. Relationship Manager

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Remote or Hybrid
United States
43000 Employees
155K-190K Annually

MetLife Logo MetLife

Customer Care Advocate AMS - Virtual - 9.28.26 - 19200

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Remote or Hybrid
United States
43000 Employees
42K-42K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account