Open Science Data Steward

Posted 17 Hours Ago
Be an Early Applicant
Emeryville, CA, USA
In-Office
100K-150K Annually
Mid level
Artificial Intelligence • Machine Learning • Biotech • Generative AI
The Role
Oversee the full lifecycle of research data, including quality, governance, metadata, publication, archival, and FAIR compliance. Partner with multidisciplinary researchers, develop documentation and scalable data infrastructure, support knowledge graph and AI/ML integration, and provide training. The role also requires navigating responsible data-sharing decisions involving privacy, safety, sensitivity, and dual-use risks while collaborating with software engineers and leadership.
Summary Generated by Built In
About Astera

Astera is a private foundation on a mission to steer science and technology toward an abundant future for all. We believe the coming years will bring an era of unprecedented scientific and technological advancement as exponential progress in AI converges with central advances in other fields to dramatically accelerate innovation. This inflection point provides an unparalleled opportunity to fundamentally rethink the institutions, systems, and tools that drive scientific progress. We are searching for visionary leaders who are compelled by the challenge of driving transformative change in the way science is funded, conducted, and communicated. You can read more about our mission, vision, and programming here.

Position Summary

We are seeking a detail-oriented and proactive Open Science Data Steward to oversee the full lifecycle of research data across our science programs. In this in-person role, you will partner with researchers from many different fields of science, maintain high data quality standards, develop scalable data infrastructure, and champion open science and FAIR data principles. The ideal candidate brings expertise in data management, a collaborative spirit, and a commitment to responsible, innovative data sharing.

Key Responsibilities
  • Partner with researchers across Astera's Science to understand the context and characteristics of their data, ensuring it is prepared and published to meet the highest standards of quality and reusability for both human researchers and AI systems.

  • Serve as one of the organization’s data stewards, overseeing the complete lifecycle of research data from creation through publication and archival, ensuring data integrity and accessibility throughout.

  • Work closely with a small but growing team of data stewards to create and maintain comprehensive living documentation for data repository interactions, establish best practices for data release timing and methods, metadata tools and standards, and ensure compliance with open science principles and FAIR (Findable, Accessible, Interoperable, Reusable) data standards.

  • Proactively identify patterns and emerging challenges in data management across research teams, collaborating with leadership to develop scalable solutions and infrastructure improvements that benefit the entire organization.

  • Support the integration of research data with open science knowledge graphs and ensure compatibility with emerging AI and machine learning ecosystems that rely on high-quality scientific data.

  • Provide training and guidance to new researchers and residents on data management best practices, fostering a culture of early and effective data sharing that maximizes scientific impact.

  • Collaborate across the organization, especially with software engineers, to identify and address technical gaps in data infrastructure, ensuring researchers have the tools and support needed to manage their data effectively.

  • Work directly with researchers and residents to ensure data policies are followed, being present on-site to navigate both routine compliance and the inevitable edge cases that require creative problem-solving, while knowing when to loop in leadership for guidance on particularly complex or sensitive data situations.

  • Navigate the balance between Astera's commitment to radically open science and the practical realities of responsible data sharing, developing thoughtful approaches for situations where research sensitivity, privacy requirements, safety considerations, or dual-use risks demand more nuanced data release strategies, and knowing when to escalate particularly challenging cases to leadership while maintaining the organization's open science mission.

Qualifications

Required:
  • Bachelor’s degree in Computer Science, Information Science, Library Science, Data Science, or related field; advanced degree preferred.

  • 3-5+ years of experience managing data in research organizations, academic institutions, or similar environments with complex data governance requirements.

  • Deep understanding of data quality principles, data governance frameworks, and metadata standards, with demonstrated ability to implement these concepts in practice.

  • Familiarity with scientific data formats, repositories, and standards.

  • Excellent communication and interpersonal skills with the ability to translate between technical data requirements and researcher needs.

  • Strong organizational skills with meticulous attention to detail and the ability to manage multiple data projects simultaneously while maintaining high standards.

  • High agency self-starter with the ability to work independently, identify opportunities for process improvement, and drive initiatives forward without constant oversight.

Preferred:
  • Experience with data across many different fields of science, including, but not limited to, life sciences, materials science, neuroscience, robotics, and atmospheric science, and the ability to adapt quickly to understand domain-specific data challenges.

  • Familiarity with the AI/ML data landscape and understanding of how to structure data for machine learning applications.

  • Knowledge of open science initiatives and experience with open data publishing platforms.

  • Experience with data integration, knowledge graphs, or semantic web technologies.

  • Track record of developing data management policies, procedures, or training materials.

Location

This role is in-person at our office in Emeryville, CA.


Astera Institute is an Equal Opportunity Employer committed to building a diverse and inclusive team.

Skills Required

  • Bachelor’s degree in Computer Science, Information Science, Library Science, Data Science, or a related field
  • 3-5+ years of experience managing data in research organizations, academic institutions, or similar environments with complex data governance requirements
  • Deep understanding of data quality principles, data governance frameworks, and metadata standards, with demonstrated ability to implement them
  • Familiarity with scientific data formats, repositories, and standards
  • Excellent communication and interpersonal skills, including translating technical data requirements for researchers
  • Strong organizational skills, meticulous attention to detail, and ability to manage multiple data projects simultaneously
  • High agency and ability to work independently, identify process improvements, and drive initiatives forward
  • Advanced degree
  • Experience working with data across multiple scientific fields
  • Familiarity with the AI/ML data landscape and structuring data for machine learning applications
  • Knowledge of open science initiatives and open data publishing platforms
  • Experience with data integration, knowledge graphs, or semantic web technologies
  • Experience developing data management policies, procedures, or training materials
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Rouen
57 Employees
Year Founded: 2020

What We Do

Astera is a private foundation with a $2.5B endowment focused on steering science and technology toward an abundant future for all. They operate like a high-velocity startup, integrating neuroscience, AI, and bioengineering for AGI research.

Similar Jobs

Capital One Logo Capital One

Artificial Intelligence Engineer

Fintech • Machine Learning • Payments • Software • Financial Services
Hybrid
4 Locations
55000 Employees
230K-286K Annually

Capital One Logo Capital One

Artificial Intelligence Engineer

Fintech • Machine Learning • Payments • Software • Financial Services
Hybrid
4 Locations
55000 Employees
230K-286K Annually

Optum Logo Optum

Service Desk Analyst

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Irvine, CA, USA
160000 Employees
24-43 Hourly

Optum Logo Optum

Principal Data Scientist

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Irvine, CA, USA
160000 Employees
227K-315K Annually

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account