Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together.
We are seeking a Senior Software Engineer to join the Enterprise Logging team and lead development of a next-generation log processing pipeline. This role is responsible for designing, building, and operating a high-throughput Lambda-based consumer that transforms and delivers ~500 TB/day of cloud log data from AWS, GCP, and Azure into Splunk Cloud. The ideal candidate has deep experience with streaming data systems, multi-cloud environments, and building production-grade data pipelines on a scale.
Primary Responsibilities:
- Design and develop a serverless event processing pipeline (AWS Lambda) that reads from Kinesis Data Streams, transforms events, and delivers CIM-compliant payloads to Splunk Cloud via HTTP Event Collector (HEC)
- Implement event parsing and transformation logic across multiple cloud log formats (CloudTrail, VPC Flow Logs, GCP Audit Logs, Azure Activity Logs, and others)
- Build and maintain source/source type mapping configurations supporting 59,000+ unique log type combinations
- Develop tiered data routing logic (Keep / Reduce / Skip) to optimize Splunk licensing costs
- Implement timestamp extraction and normalization across heterogeneous event formats
- Operate and monitor the pipeline in production - build observability dashboards, alerting, and error handling (dead-letter queues, retry logic)
- Collaborate with the Centralized Logging team on Kinesis stream integration and Enhanced Fan-Out configuration
- Participate in architecture decisions, code reviews, and CI/CD pipeline development
- Write infrastructure as code (Terraform or CloudFormation) for all deployed components
- Contribute data standards and onboarding documentation for new log sources
- Comply with all applicable Company policies, procedures, and business directives, changes including those relating to work location, team assignments, work schedules, and flexible work arrangements
Required Qualifications:
- Cloud Platforms (Multi-Cloud)
- AWS (primary): Solid hands-on experience with Lambda, Kinesis Data Streams, S3, IAM, CloudWatch, PrivateLink, DynamoDB, and SQS
- GCP: Working knowledge of Cloud Logging, Pub/Sub, IAM, and GCP audit log formats
- Azure: Working knowledge of Event Hubs, Activity Logs, Diagnostic Logs, Storage Accounts, and Azure Monitor
- Streaming & Data Engineering
- Experience building high-throughput streaming pipelines (Kinesis, Kafka, or equivalent)
- Experience with data serialization formats (JSON, AVRO, Protobuf) and compression (gzip, base64 encoding)
- Understanding of event-driven architecture and at-least-once delivery semantics
- Familiarity with fan-out patterns, backpressure handling, and shard-level parallelism
- Software Development
- 5+ years of professional software engineering experience
- Experience with CI/CD pipelines (GitHub Actions, Jenkins, or equivalent)
- Experience writing and maintaining infrastructure as code (Terraform preferred; CloudFormation acceptable)
- Proficiency in Go (primary) with experience writing and maintaining production Lambda functions; the existing consumer is implemented in Go
- Solid understanding of error handling patterns: exponential backoff, dead-letter queues, circuit breakers
- Proficiency in Python (secondary) for tooling, automation, and supporting services
- Proficiency with Git-based workflows, code review practices, and trunk-based development
- Observability & Operations
- Experience building monitoring and alerting data pipelines for production
- Understanding of distributed systems failure modes and operational runbook development
- Familiarity with CloudWatch Metrics, Logs, and Alarms (or equivalent observability stack)
What You'll Be Building
This is not a maintenance role. You will be building a greenfield, high-scale data pipeline that:
- Processes ~500 TB/day across three major cloud providers
- Transforms events from a legacy AVRO-wrapped format into CIM-compliant native payloads
- Directly enables Splunk Enterprise Security detection capabilities (correlation searches, data model acceleration)
- Implements intelligent data tiering to reduce licensing costs by tens of terabytes per day
At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone-of every race, gender, sexuality, age, location and income-deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes - an enterprise priority reflected in our mission.
#NJP#IJP
Skills Required
- 5+ years of professional software engineering experience
- Hands-on AWS experience with Lambda, Kinesis Data Streams, S3, IAM, CloudWatch, PrivateLink, DynamoDB, and SQS
- Working knowledge of GCP Cloud Logging, Pub/Sub, IAM, and GCP audit log formats
- Working knowledge of Azure Event Hubs, Activity Logs, Diagnostic Logs, Storage Accounts, and Azure Monitor
- Experience building high-throughput streaming pipelines using Kinesis, Kafka, or equivalent
- Experience with JSON, Avro, Protobuf, gzip, and Base64 encoding
- Understanding of event-driven architecture and at-least-once delivery semantics
- Familiarity with fan-out patterns, backpressure handling, and shard-level parallelism
- Proficiency in Go for production Lambda functions
- Proficiency in Python for tooling, automation, and supporting services
- Experience with CI/CD pipelines such as GitHub Actions or Jenkins
- Experience writing and maintaining infrastructure as code, preferably Terraform or CloudFormation
- Understanding of exponential backoff, dead-letter queues, and circuit breakers
- Proficiency with Git workflows, code reviews, and trunk-based development
- Experience building monitoring and alerting for production data pipelines
- Understanding of distributed-systems failure modes and operational runbook development
- Familiarity with CloudWatch Metrics, Logs, and Alarms or equivalent observability tools
What We Do
Optum, part of the UnitedHealth Group family of businesses, is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together. At Optum, we support your well-being with an understanding team, extensive benefits and rewarding opportunities. By joining us, you’ll have the resources to drive system transformation while we help you take care of your future. We recognize the power of connection to drive change, improve efficiency and make a difference in health care. Join a team where your skills and ideas can make an impact and where collaboration is key to creating technology that produces healthier outcomes.
Gallery
Optum Offices
Hybrid Workspace
Employees engage in a combination of remote and on-site work.
Optum has three workplace models that balance the needs of the business and the responsibilities of each role. These models, core on‑site (5 days/week), hybrid (4 days/week) and telecommute or fully remote, vary by country, role and location.