Data Engineer Jobs in Sunnyvale, CA

One Month AgoSaved
Remote or Hybrid
Santa Clara, CA, USA
191K-334K Annually
Senior level
191K-334K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Build and operate a security data platform connecting SIEM, identity, asset, threat intelligence, and knowledge graph systems. Design core APIs, distributed data pipelines, schemas, risk scoring, entity resolution, and coverage queries. Own architecture and platform workstreams from scoping through delivery, collaborate with detection and identity teams, review code, mentor engineers, and drive partner adoption. The role requires expert Python, deep Splunk expertise, cybersecurity experience, and senior-level project ownership.
Top Skills: Asset InventoryAsynchronous Data PipelinesIdentity GovernanceKnowledge GraphsLarge Language ModelsMachine LearningMessage BusMitre Att&CkPythonRetrieval-Augmented GenerationSchema RegistrySIEMSplSplunkSplunk Apis
2 Days AgoSaved
Hybrid
Sunnyvale, CA, USA
117K-234K Annually
Senior level
117K-234K Annually
Senior level
Big Data • Cloud • Logistics • Machine Learning • Retail
Designs and deploys scalable data solutions supporting AI agents, LLM workflows, and autonomous decision-making. Builds cloud-based pipelines, semantic layers, data models, monitoring, governance, and full-stack services. Partners with engineering, AI/ML, product, and business teams to deliver reliable, traceable, high-performing data products and enterprise applications. Requires expertise in distributed data processing, cloud platforms, databases, APIs, DevOps, and modern frontend and backend development.
Top Skills: AngularAzureAzure SqlBigQueryCassandraCi/CdCosmosCrewaiDataflowDockerDruidGCPJavaKafkaKubernetesLangchainLlamaindexMongoDBPrestoPub/SubPysparkPythonReactScalaSparkSpark StreamingSpring BootSQLVue
2 Days AgoSaved
Hybrid
Sunnyvale, CA, USA
143K-286K Annually
Senior level
143K-286K Annually
Senior level
Big Data • Cloud • Logistics • Machine Learning • Retail
Designs and operates billion-scale data platforms supporting streaming, batch, advertising, marketing, measurement, and machine learning workloads. Builds reliable, tenant-aware pipelines with replayability, fault recovery, observability, quality controls, and cost efficiency. Leads domain architecture, performance optimization, design reviews, mentoring, incident response, and governed automation across global advertising products.
Top Skills: Apache AirflowApache KafkaSparkBigQueryCi/CdDataflowDataprocETLGoogle Cloud Platform (Gcp)Google Cloud Storage (Gcs)JavaPub/SubPythonScalaSQL
Reposted 2 Days AgoSaved
In-Office or Remote
5 Locations
152K-288K Annually
Senior level
152K-288K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Design and implement high-performance, scalable storage systems and client libraries for AI workloads. Build data loading, checkpointing, caching, POSIX-style and object-store integrations, observability and telemetry, and validate performance with platform, SRE, and AI teams using modern engineering practices.
Top Skills: C/C++CachingCloud InfrastructureFilesystemsGoJavaKubernetesLinuxObject StoresPosixPythonRust
Reposted 2 Days AgoSaved
In-Office or Remote
Santa Clara, CA, USA
152K-288K Annually
Senior level
152K-288K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Design and implement cloud-native data management services (catalog, metadata, dataset/checkpoint lifecycle) for exabyte-scale GPU workloads. Build backend systems on Kubernetes and major cloud providers, collaborate with research and cross-functional teams, document architecture, and drive integration with storage/compute innovations (GPU Direct Storage, DPU).
Top Skills: Apache IcebergSparkAWSAzureC/C++Cloud NativeData LakeDpuFeature StoresGCPGoGpu Direct StorageJavaKubernetesMetadata ManagementObject StoragePython
3 Days AgoSaved
In-Office
4 Locations
120K-261K Annually
Senior level
120K-261K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Design, build, secure, and operate Microsoft’s clean room data platform and infrastructure. Develop scalable data onboarding, identity resolution, audience activation, measurement, and partner collaboration capabilities. Implement access controls, encryption, network isolation, monitoring, auditing, compliance, and reliability standards for sensitive data environments. Automate provisioning, deployments, and operational controls while leading incident investigations and platform improvements. Partner with engineering, privacy, security, marketing, data engineering, and external organizations on long-term privacy-safe data collaboration solutions.
Top Skills: AlertingSparkAudience MatchingAzureAzure Data FactoryAzure Event HubsAzure SynapseCi/CdConfidential ComputingConsent ManagementCustomer Data PlatformsDatabricksEncryptionIdentity And Access ManagementIdentity GraphsIdentity ResolutionInfrastructure As CodeLoggingMicrosoft FabricMonitoringObservability
4 Days AgoSaved
In-Office
Mountain View, CA, USA
70-70 Annually
Internship
70-70 Annually
Internship
Automotive
Develop large-scale data pipelines, infrastructure, and tools for generating, analyzing, and evaluating machine learning datasets from autonomous driving logs and simulation data. Build robust data platforms supporting Waymo Driver launch evaluations, collaborate with Data Science and Quantitative Analytics teams, and improve dataset curation, sampling, slicing, accuracy, and completeness.
Top Skills: C++Data InfrastructureData PipelinesDistributed SystemsMachine LearningPythonSQL
Reposted 8 Days AgoSaved
In-Office or Remote
Santa Clara, CA, USA
184K-357K Annually
Senior level
184K-357K Annually
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Designs and builds scalable microservices, distributed applications, ETL pipelines, and AI data workflows for autonomous vehicle data. Responsibilities include streaming, orchestration, security, data management, model deployment, RAG workflows, video curation, behavioral search, and optimization of large-scale distributed systems. The role collaborates with software and AI experts while using technologies such as Python or Golang, Kubernetes, Hive, Parquet, SQL, vector databases, Spark, and NVIDIA RAPIDS.
Top Skills: Apache HiveApache ParquetSparkData PipelinesDistributed ComputingETLGoHelmKubernetesLarge Language Models (Llms)MicroservicesMilvusNvidia RapidsPythonRetrieval-Augmented Generation (Rag)SQLVector DatabasesVision-Language Models (Vlms)
One Month AgoSaved
Remote or Hybrid
Sunnyvale, CA, USA
275K-348K Annually
Expert/Leader
275K-348K Annually
Expert/Leader
Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
Leads an engineering team building scalable ML data infrastructure, including data ingestion, processing, storage, access, pipelines, and analytical platforms. Defines architecture, evaluates technologies, improves performance and reliability, translates data requirements into technical solutions, and ensures governance, security, and regulatory compliance. Provides mentorship, strategic roadmap input, and cross-functional collaboration with ML engineers, data scientists, platform engineers, and product teams.
Top Skills: Agile MethodologyAnalytical PlatformsData EngineeringData InfrastructureData PipelinesJavaMachine LearningPythonScala
One Month AgoSaved
In-Office or Remote
8 Locations
153K-270K Annually
Senior level
153K-270K Annually
Senior level
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
Build and operate Block’s data platform, including warehousing, orchestration, business intelligence, governance, and agentic workflows. Develop reliable Python services, developer tools, and Terraform infrastructure for petabyte-scale data. Partner with data practitioners and product teams, establish engineering standards, provide technical direction, and mentor engineers while ensuring security, reliability, observability, usability, and data governance.
Top Skills: AnomaloApache AirflowApache IcebergAWSBuildkiteDatabricksDatahubDbtGitopsKotlinLookerOmniPrefectPythonSnowflakeTerraform
One Month AgoSaved
Remote or Hybrid
8 Locations
153K-270K Annually
Senior level
153K-270K Annually
Senior level
Blockchain • Fintech • Mobile • Payments • Software • Financial Services
Build and operate a scalable data platform supporting warehousing, orchestration, business intelligence, governance, and agentic workflows. Develop Python services, Terraform infrastructure, and developer tools for petabyte-scale data. Partner with data practitioners and product teams, establish security and reliability patterns, set technical direction, improve engineering standards, and mentor engineers.
Top Skills: AnomaloApache AirflowApache IcebergAWSBuildkiteDatabricksDatahubDbtGitopsKotlinLookerOmniPrefectPythonSnowflakeTerraform
16 Days AgoSaved
Hybrid
Santa Clara, CA, USA
200K-250K Annually
Senior level
200K-250K Annually
Senior level
Artificial Intelligence • Machine Learning • Software • Analytics
Leads hands-on design, development, and optimization of AI data movement and distributed storage systems. Responsibilities include integrating NVIDIA NIXL and DDN Infinia with GPU inference platforms, optimizing GPU-to-storage I/O using GPUDirect Storage, RDMA, and NVMe-over-Fabrics, developing KV cache and multi-tier storage strategies, benchmarking production systems, resolving performance bottlenecks, influencing distributed inference architecture, and mentoring engineers.
Top Skills: CC++Ddn InfiniaDistributed StorageGpu ComputingHpcInfinibandKv Cache ManagementLinuxLlm InferenceNvidia Gpudirect StorageNvidia NixlNvmeNvme-Over-FabricsObject StoragePythonPyTorchRdmaRetrieval-Augmented GenerationScalable File SystemsSsdTensorFlowVector Databases
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted 16 Days AgoSaved
In-Office
3 Locations
141K-250K Annually
Senior level
141K-250K Annually
Senior level
eCommerce • Legal Tech
Architects, builds, and scales data pipelines, ETL processes, data models, and platforms using dbt, Snowflake, Airflow, and Fivetran. Leads data engineering best practices including CI/CD, data quality, governance, and observability. Collaborates with Data Science, Engineering, Product, and business stakeholders; mentors engineers and data scientists; contributes to architecture and technical decisions; and participates in on-call support.
Top Skills: Amazon RedshiftApache AirflowCi/CdDbtFivetranGoogle BigqueryPythonSnowflakeSQL
17 Days AgoSaved
In-Office
Santa Clara, CA, USA
168K-311K Annually
Expert/Leader
168K-311K Annually
Expert/Leader
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Develop and manage enterprise data platforms, pipelines, APIs, MDM/RDM solutions, governance frameworks, and AI-enabled observability for semiconductor supply chain data. Integrate SAP, PLM, planning, and business systems while building canonical models, metadata catalogs, lineage, data quality processes, and agentic AI workflows. Partner with engineering, business, and IT stakeholders to deliver scalable, governed data solutions across manufacturing and supply chain operations.
Top Skills: AnaplanSparkAPIsChatgptClaudeCopilotDatabricksDelta LakeEltETLEvent-Driven ArchitectureGeminiInformatica CaiInformatica CdgcInformatica CdiInformatica Data CatalogInformatica IdqInformatica Intelligent Data Management Cloud (Idmc)Informatica MdmInformatica Metadata ManagementInformatica Reference 360Knowledge GraphsLlmsMicroservicesOntologyPysparkRagSalesforce (Sfdc)Sap IbpSap MdgSap S/4Hana
Reposted 18 Days AgoSaved
In-Office
Sunnyvale, CA, USA
200K-275K Annually
Senior level
200K-275K Annually
Senior level
Hardware • Industrial
Build and optimize edge and cloud data pipelines for perception ML models. Scale data engine for high-fidelity labels, reduce annotation costs, integrate foundation models for automated labeling and QA, leverage software/hardware-in-the-loop testing, and support DoD field use cases and model deployment lifecycle.
Top Skills: DatabasesDockerGoGpusHardware-In-The-LoopKubernetesLlmsMicroservicesMlopsMultimodal ModelsOpensearchPostgresPythonReactSoftware-In-The-LoopTypescriptVlms
20 Days AgoSaved
In-Office
Mountain View, CA, USA
Senior level
Senior level
eCommerce • Fintech • Logistics • Retail
Build and manage high-volume analytics and production ETL/ELT systems. Partner with Data Science, Engineering, and Product teams to ensure reliable Search data logging and usage. Own data quality, observability, security, compliance, SLAs, scalability, performance, and cost efficiency across the data lifecycle. The role requires expertise in Python, SQL, data modeling, distributed data technologies, massive datasets, and data pipeline optimization.
Top Skills: AirflowAmazon S3DagsterData ModelingDbtEltETLHdfsPythonSQL
22 Days AgoSaved
In-Office or Remote
9 Locations
125K-220K Annually
Mid level
125K-220K Annually
Mid level
Aerospace • Hardware • Software
Build and operate reliable data pipelines transforming flight test, simulation, and operational data into versioned, traceable ML training datasets. Responsibilities include ingestion, normalization, time alignment, validation, dataset lineage, anomaly mining, labeling workflows, deployment feedback loops, monitoring, and data coverage analysis. Collaborate with flight test and operations teams on instrumentation and data access. The role is on-site in Boston and requires production data engineering experience supporting machine learning pipelines.
Top Skills: PythonSQL
24 Days AgoSaved
Remote or Hybrid
7 Locations
130K-165K Annually
Senior level
130K-165K Annually
Senior level
Retail • Analytics
Own production-grade analytics pipelines from design through maintenance, build dbt-based analytical models using Kimball dimensional modeling, and deliver reliable data features. Partner with Product, Engineering, and Data Science teams to translate retail signals into business insights. Establish standards for monitoring, documentation, reproducibility, and code quality. Use AI tools to accelerate development and prototype conversational data experiences for natural-language retail data exploration.
Top Skills: Ai AgentsAirflowBigQueryCi/CdDbtGCPKimball Dimensional ModelingLlm ApisLookerSQLTableau
24 Days AgoSaved
Remote or Hybrid
7 Locations
135K-170K Annually
Senior level
135K-170K Annually
Senior level
Retail • Analytics
Own and scale the company’s analytics infrastructure, including BigQuery data warehouses, dbt data marts, external data exchange, DataOps, data quality, governance, security, and cloud cost optimization. Partner with application database teams, analytics engineers, and data scientists to enable reliable metrics and scalable client data delivery.
Top Skills: AirflowApache BeamBigQueryDbtGoogle Cloud PlatformGoogle Cloud StorageJIRAPostgresPythonSQLTerraform
Reposted One Month AgoSaved
In-Office
Mountain View, CA, USA
166K-244K Annually
Mid level
166K-244K Annually
Mid level
Artificial Intelligence • Greentech • Hardware • Internet of Things • Transportation • Cybersecurity • Automation
Build and consolidate data infrastructure for ML training: design automated ETL/ELT pipelines, implement DataOps practices (validation, monitoring, anomaly detection), integrate annotation workflows, manage dataset versioning and storage, and collaborate with ML engineers and operations to produce high-quality, reproducible training datasets.
Top Skills: Apache AirflowBigQueryCloud StorageDagsterData ValidationDataflowDataprocDvcGoogle Cloud ComposerGreat ExpectationsNumpyPandasPrefectPythonSQLTfxVertex Ai Data Pipelines
Reposted One Month AgoSaved
In-Office
Mountain View, CA, USA
120K-180K Annually
Senior level
120K-180K Annually
Senior level
Big Data • eCommerce • Robotics
The Data Engineer will analyze data quality, design prediction algorithms, collaborate with the engineering team, and generate business insights.
Top Skills: CassandraHadoopPythonSparkSQLTableau
Reposted One Month AgoSaved
In-Office
2 Locations
165K-242K Annually
Senior level
165K-242K Annually
Senior level
Cloud • Information Technology • Machine Learning
Design and implement data platforms, manage data infrastructure, develop streaming architectures, improve performance, and ensure data compliance.
Top Skills: CockroachdbGoJavaKafkaKubernetesNatsPythonTidbYdbYugabyte
Reposted One Month AgoSaved
In-Office
5 Locations
207K-275K Annually
Expert/Leader
207K-275K Annually
Expert/Leader
Cloud • Information Technology • Machine Learning
Lead architecture and standards for CoreWeave's enterprise data ecosystem: design lakehouse/streamhouse architectures, modeling and semantic layers, governance, data quality, lineage, observability, and reusable data products. Drive cross-domain solutions, resolve performance and scalability constraints, evaluate core data technologies, conduct architecture reviews, and mentor senior engineers.
Top Skills: Apache FlussApache HudiApache IcebergApache PaimonAutomqClickhouseData VaultDelta LakeFlinkJavaKafkaKubernetesPulsarPythonRustScalaSparkSQLStarrocksTrino
Reposted One Month AgoSaved
In-Office
Mountain View, CA, USA
155K-210K Annually
Mid level
155K-210K Annually
Mid level
Automotive
Hybrid hardware/software engineer responsible for characterizing 4D LiDAR sensors, supporting customer integrations, building production Python/C++ data pipelines, running calibration and field debugging, maintaining ETL for high-volume sensor data, and coordinating data labeling and collection operations.
Top Skills: Binary Data FormatsC++Camera SystemsEtl PipelinesGpsLidarPoint Cloud AnalysisPythonRadarRosSensor Logs
28 Days AgoSaved
In-Office
Mountain View, CA, USA
162K-234K Annually
Senior level
162K-234K Annually
Senior level
Automotive • Software
Own the architecture and operation of large-scale autonomous-driving datasets, from multimodal vehicle-data ingestion through curation, quality monitoring, labeling, versioning, and ML training delivery. Design schemas, storage, query, distributed pipelines, observability, and reproducible dataset splits. Lead data-quality, drift, anomaly, auto-labeling, human-review, and closed-loop workflows, including trajectory preparation for imitation learning and offline reinforcement learning. Provide staff-level technical leadership across autonomy, annotation, storage, governance, and model teams.
Top Skills: Amazon S3Apache AirflowApache ArrowApache BeamSparkDagsterDaskGoogle Cloud StorageLanceParquetPythonRaySQLVlms
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account