Maximum of 25 job preferences reached.
Top Machine Learning Engineer Jobs
Artificial Intelligence • Automotive • Information Technology • Robotics
Lead ML efforts for behavior and planning: design, train, and deploy state-of-the-art models (prediction, decision-making, RL/imitation, generative, transformers), own full model lifecycle from data to onboard inference, generalize behavior across geographies and platforms, collaborate cross-functionally, and mentor engineers.
Top Skills:
Python
Food
Lead MLOps efforts to productionize machine learning models: build containerized deployments, automate CI/CD pipelines, manage Azure-based ML lifecycle, monitor model performance and data drift, collaborate with data engineering on pipelines, and mentor junior MLOps engineers while driving engineering best practices.
Top Skills:
Aws SagemakerAzure Container Instances (Aci)Azure Container RegistryAzure DevopsAzure Kubernetes Service (Aks)Azure Machine LearningCi/CdContainerizationDockerGcp Vertex AiGitGithub ActionsKubernetesMedallion ArchitectureMicrosoft Fabric Data SciencePython
Artificial Intelligence • Machine Learning • Natural Language Processing • Software
Design, build, and own ML evaluation systems and production-grade NLP/LLM solutions. Translate research into scalable agentic AI systems, collaborate with Research/Product/Platform, ensure quality, safety, and performance, and mentor engineers.
Top Skills:
AthenaAWSCi/CdDockerKafkaKubernetesLlmNlpPython
Artificial Intelligence • Big Data • Machine Learning
Lead hands-on research and engineering across agentic ML: train and fine-tune models, design evaluation and observability, build production improvement loops and tooling, prototype agent architectures, partner cross-functionally to productionize research, and set technical direction while mentoring senior engineers and scientists.
Top Skills:
Agent ArchitecturesContinuous Fine-TuningEvaluation And Observability InfrastructureInferenceLarge Language Models (Llms)Multi-Agent OrchestrationOnline LearningProduction Ml SystemsReinforcement LearningRetrieval/Memory SystemsReward ModelingRlaifRlhfSft (Supervised Fine-Tuning)Training Pipelines
Big Data • Software
Design, build, and productionize ML systems focused on information extraction, retrieval, ranking, and recommendation. Own full ML lifecycle from feature engineering to deployment, observability, and evaluation. Collaborate with data, infra, and product teams to scale serving for streaming and batch inference and improve models using dataset engineering and graph/embedding techniques.
Top Skills:
A/B TestingCi/CdEmbedding ModelsGraph DatabasesLearn-To-RankMl Model ServingOcrPythonPyTorchScikit-LearnVector Databases
Healthtech
Lead design and deployment of production-grade agentic AI systems for healthcare. Architect scalable multi-agent orchestration, data curation pipelines, evaluation frameworks, and translate ML research to secure, low-latency clinical applications while mentoring engineers and collaborating with clinical and product teams.
Top Skills:
Distributed Machine LearningJaxLarge Language Models (Llms)Multi-Agent OrchestrationPrompt OptimizationPythonPyTorchTensorFlow
Artificial Intelligence • Hardware • Software • Semiconductor
Design and implement system-level debugging, validation, and observability platforms. Build automated anomaly collection/analysis, visualization and root-cause tools, failure classification and monitoring frameworks. Extend compilers, runtimes and programming interfaces for profiling and instrumentation, improve bring-up and low-level debug workflows, lead cross-functional initiatives, support incident response, and establish debuggability and reliability best practices.
Top Skills:
C++CompilersFirmwareHardware InterfacesInstrumentationProfilingProgramming InterfacesPythonRuntimesVisualization Tools
Artificial Intelligence • Hardware • Software • Semiconductor
Bring up and validate next-generation AI hardware systems; debug complex hardware-software integration issues using logs, telemetry and diagnostics; build automation, testing, and tooling to improve validation, observability, and debugging workflows; collaborate with hardware teams to reproduce, triage, and resolve failures as systems move toward production.
Top Skills:
C++Distributed SystemsIpcLinuxLog AnalysisNetworkingOperating SystemsPerformance AnalysisPythonSystem Telemetry
Artificial Intelligence • Hardware • Software • Semiconductor
Drive end-to-end ML model inference performance: build kernel- and system-level performance models, optimize kernel microcode and compiler algorithms, debug runtime performance on system and cluster, and develop tooling to visualize and analyze performance data from the Wafer Scale Engine and compute cluster.
Top Skills:
C++Cerebras Wafer Scale EngineCompilersCpu/Gpu SimulatorsHpcKernel MicrocodePerformance Profiling ToolsPython
Artificial Intelligence • HR Tech • Software • Generative AI
Design and build reinforcement-learning training environments and diverse tasks to evaluate and improve LLM agents; iterate rapidly on task designs from customer feedback, deliver high-quality outputs with minimal supervision, and maintain PST overlap for collaboration.
Top Skills:
Large Language Models (Llms)Machine LearningPythonReinforcement Learning
Hardware
Develop and maintain high-performance C++ machine control and data analysis software for advanced mask inspection systems. Integrate AI/ML capabilities, collaborate with multidisciplinary engineering teams, optimize and scale existing code, and support integration/testing and customer escalations. Work with APIs, containers, and observability tools to enable reliable production deployment.
Top Skills:
AIAngularC++DockerGrafanaGtkKubernetesLinuxMachine LearningPostgresPrometheusQtReactRestRpcSingularityVue
Artificial Intelligence • Healthtech • Insurance • Software
Design, build, and deploy production ML systems that combine structured clinical/EHR data with computer vision outputs. Own model calibration, interpretability, monitoring, drift detection, retraining workflows, and regulatory documentation. Partner cross-functionally, mentor engineers, and translate clinical questions into reliable, production-grade ML solutions.
Top Skills:
Ci/CdComputer VisionDrift DetectionEhrLongitudinal ModelingMachine Learning PipelinesModel MonitoringMultimodal ModelsPythonStructured Clinical DataSurvival AnalysisTestingTime-Series AnalysisTraining And Inference Infrastructure
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
News + Entertainment
Design, build, deploy, and operate production-grade ML systems and pipelines for a massive consumer audience. Own full lifecycle from data processing and feature engineering through training, validation, deployment, and monitoring. Drive ML infrastructure and architecture decisions, partner with product and data teams, optimize models against business metrics, and mentor engineers to ensure security, reliability, and compliance.
Top Skills:
Automated Model Lifecycle ManagementCloud-Native Ml ArchitecturesDistributed TrainingMlops PlatformsModel Monitoring/ObservabilityPython
Aerospace • Hardware • Software • Defense • Manufacturing
Build and operate Hadrian's production ML platform: standardize MLflow/Dagster-based deployments, serve batch and online inference, maintain feature serving and lineage, detect drift, enable model CI/CD, observability, autoscaling, and developer tooling for reliable model production across factories.
Top Skills:
Amazon EcrAmazon EksBentomlContainersDagsterFastapiFeastGoGpu InferenceGrpcKserveKubernetesMlflowPythonRay ServeRustSagemakerSQLTectonTritonVertex Ai
Artificial Intelligence • Software
The Research Engineer, Infrastructure will build distributed training systems, optimize performance, manage data pipelines, and enhance research workflows, ensuring the infrastructure scales with AI advancements.
Top Skills:
C++GpuPythonPyTorch
Cloud • Security • Software • Cybersecurity
Build and operate model validation, quantization, and safety systems for production ML. Develop pipelines for security scanning, optimization (quantization/pruning), routing, prompt management, and evaluation frameworks measuring accuracy, performance, and safety across model lifecycles.
Top Skills:
AwqCi/CdCloud InfrastructureContainerizationGgufGptqJaxLlmsPythonPyTorchTensorFlowTransformers
Retail
Develop, deploy, and support secure, scalable production software and AI/ML solutions. Collaborate with product, UX, and engineering to deliver features, create tests and automation, instrument monitoring and dashboards, and participate in agile processes while continuously learning and improving team practices.
Top Skills:
APIsAWSAzureAzure AiCi/CdCSSDevOpsGCPGitHTMLJavaJavaScriptLangchainMicroservicesMlopsNoSQLOpenaiPrompt EngineeringPythonPyTorchRelational DatabasesScikit-LearnSQLTensorFlowTypescript
Information Technology • Software • Analytics • Cybersecurity
Design, build, and maintain scalable MLOps pipelines and data workflows to train, fine-tune, deploy, and monitor multilingual text and speech translation models. Prepare and optimize multilingual datasets, support model evaluation and benchmarking, and enable low-resource language training while collaborating with data scientists to integrate models into production.
Top Skills:
BedrockDockerEc2IamLambdaPythonS3SagemakerVpc
Software
Design, build, and operate core AI capabilities and reusable model integrations for an enterprise platform. Integrate foundation models into production, optimize performance and reliability, prototype new techniques, and support scalable backend services and distributed AI systems in collaboration with cross-functional teams.
Top Skills:
Ai AgentsAPIsCloud-Native InfrastructureDistributed SystemsFoundation ModelsKubernetesLarge Language ModelsModel OrchestrationMultimodal AiObservabilityPythonRetrieval-Augmented Generation (Rag)Rust
Fintech • Payments
Design, deploy, and productize scalable traditional and generative AI/ML solutions. Collaborate with data scientists and engineers, build LLM-based agents, ensure code quality and reliable CI/CD deployments, mentor team members, and drive adoption of state-of-the-art ML techniques.
Top Skills:
Agent DevelopmentSparkAWSAzureCi/CdDockerGCPJavaKubernetesLarge Language ModelsMapreducePython
Automotive • Internet of Things • Mobile • Semiconductor • Industrial
Develop and prototype AI/ML-driven automation for ASIC/SoC design and implementation. Collaborate with academia and internal teams to apply GenAI, RL, GNNs, and CNNs to VLSI CAD flows, pilot solutions on real designs, and improve power, performance, area, and quality.
Top Skills:
BitbucketC++Ci/CdConvolutional Neural NetworksDesignsyncEda ToolsGenaiGitGraph Neural NetworksJenkinsLlmsLsfMakePythonRecurrent Neural NetworksReinforcement LearningSplunkUnix/LinuxVlsi Cad
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
As a Senior Software Engineer at NVIDIA, you will design and implement inference software optimizations for AI applications using TensorRT, collaborating with deep learning experts to influence hardware and software design.
Top Skills:
C++CudaPython
Aerospace • Energy
This role involves designing, developing, and maintaining AI/ML products, collaborating with stakeholders, and ensuring model deployment on AWS. Responsibilities include creating APIs, establishing MLOps practices, and driving AI strategy for operational improvements.
Top Skills:
AWSC#DatabricksFastapiFlaskJavaMlflowPythonTypescript
Reposted 9 Days AgoSaved
3D Printing • Consulting • Design • Manufacturing
Build scalable agentic frameworks and reproducible experimental pipelines that integrate frontier multimodal models with rich neuroscience data. Implement orchestration, MLOps, and cloud infrastructure, balancing research and production engineering on a small collaborative team.
Top Skills:
Agent Orchestration Frameworks (Claude Code Agent TeamsAgentic FrameworksAzure)BeadsClaude Computer UseCodexCodex Multi-AgentsCrew Ai Agent Teams)Fdm-1GCPGemini-CliManusMicrosoft Agent FrameworkMlflowMultimodal Large Language ModelsOai OperatorOpenspecOrchestration Platforms (AwsPrompt EngineeringPythonPyTorchQwen Code)TransformersTui/Cli Coding Agents (Claude CodeWeights & Biases
Reposted 9 Days AgoSaved
3D Printing • Consulting • Design • Manufacturing
Optimize training and inference performance of foundation models by writing and tuning CUDA/Triton GPU kernels, profiling and removing bottlenecks, implementing low-precision and mixed-precision strategies, optimizing MoE routing and expert dispatch, integrating high-performance libraries, and collaborating closely with researchers to translate model ideas into scalable, efficient implementations.
Top Skills:
CudaCudnnCutlassFlashattentionGpu KernelsKv-Cache ManagementLow-Precision TrainingMixed-Precision TrainingMoe (Mixture Of Experts)Moe RoutingNcclNsight ComputeNsight SystemsPost-Training QuantizationPyTorchPytorch ProfilerQuackSpeculative DecodingTensorrt-LlmTransformer ArchitecturesTritonVllm
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Companies Hiring Machine Learning Engineers
See AllPopular Job Searches
All Filters
Total selected ()
No Results
No Results

.png)

.png)









.png)
















