Software Engineer - Research Data Platform

Posted 2 Days Ago
Be an Early Applicant
Hiring Remotely in Berlin, DEU
In-Office or Remote
Mid level
Artificial Intelligence • Biotech
We build foundation models that will transform biology.
The Role
Build and scale the research data platform powering biological AI workloads. Responsibilities include designing schemas and storage systems, optimizing distributed storage and database performance, developing typed APIs, improving pipelines and cloud efficiency, and implementing data validation, versioning, testing, and access controls. The role collaborates with researchers and product engineers to deliver maintainable, user-friendly infrastructure for scientists, engineers, and AI agents.
Summary Generated by Built In

Bioptimus is building the first universal AI foundation model for biology to fuel breakthrough discoveries and accelerate innovation in biomedicine. With more than $75M in funding, Bioptimus is a fast-growing start-up headquartered in Paris, incorporated in October 2023. Backed by leading international venture capitalists, our world-class team of scientists and engineers is redefining the frontiers of AI and life sciences. 

Software Engineer - Research Data Platform

Paris / Remote EU


Bioptimus is building the best-in-class universal AI foundation model for biology to fuel breakthrough discoveries and accelerate innovation in biomedicine. With more than $75M in funding, Bioptimus is a fast-growing startup incorporated in October 2023 and headquartered in Paris. Backed by leading international venture capitalists, our world-class team of scientists and engineers is redefining the frontiers of AI and life sciences.

This is a remote role. We’re headquartered in Paris, but the position can be performed remotely outside of Paris.

About the Role

We are a fast-moving, data-centric start-up on a mission to bridge the gap between complex biological data and cutting-edge AI. As a Software Engineer in our research data platform team, you will help develop the backbone of our data architecture, designing and scaling the systems that power our AI models and user-facing tools, both internal and external.

We are looking for someone passionate about scalable, efficient, and highly structured data storage. In particular, we are looking for someone interested in designing systems that account for the complex structures inherent in biological data. You will build clean, maintainable systems that make massive biological datasets accessible, reliable, and actionable. If you love optimizing performance, improving schemas, and seeing your work directly empower a broad audience of stakeholders—from scientists and product engineers to AI agents—you will fit right in.

This is a mid-level to senior individual contributor role. You will collaborate closely with our multidisciplinary team of researchers and engineers to drive software development and productization efforts.

What You Will Be Doing

As a Software Engineer for our research data platform, you will own the following responsibilities:

  • Architect and build for performance: Design, implement, and maintain robust, scalable data schemas and storage solutions optimized for high-performance AI workloads.
  • Optimize storage formats: Benchmark, profile, improve, and extend distributed storage using chunking, compression, parallelization, and custom solutions.
  • Produce clean code: Develop and maintain high-quality, production-ready data systems, following clean code principles and engineering best practices.
  • Design interfaces for people and AI agents: Build clean, typed, well-documented APIs that let researchers, engineers, and AI agents query and extend the platform programmatically.
  • Collaborate and drive delivery: Work with researchers and product engineers to scope needs, align on priorities, and own projects end-to-end.
  • Performance tuning: Monitor, profile, and optimize database queries, storage read and write paths, pipeline bottlenecks, and cloud infrastructure costs.
  • Data governance and security: Collaborate with platform engineers to implement rigorous data validation, testing, versioning, and access control.
What You Will Bring

The successful candidate will have a team-first attitude, be independent, curious, and detail-oriented, thrive in a dynamic, fast-paced environment, and be fun to work with. Moreover, we value individuals with the following skills:

Technical and Professional Qualifications
  • Python expertise: Deep, production-level knowledge of Python with a passion for clean, readable, and highly maintainable code.
  • Backend frameworks: Strong hands-on experience with modern Python data tools and frameworks, such as Pydantic (data validation), SQLAlchemy (ORM), Alembic (database migrations), object storage abstractions, and FastAPI or similar frameworks.
  • Structured databases: Expertise in relational database management systems (RDBMSs) such as PostgreSQL, including schema design, indexing strategies, and query optimization.
  • Interfaces and agents: Experience designing API surfaces and exposure to protocols for programmatic and agent access such as the Model Context Protocol (MCP).
  • User-centric mindset: A strong belief that data infrastructure is a product, combined with a commitment to keeping it usable and accessible to non-technical stakeholders.
How to Stand Out

Each of the following would be a valuable bonus, not a requirement:

  • Biotech/life sciences affinity: Prior experience handling biological data formats (e.g., histology, transcriptomics, genomics, proteomics, or clinical trial data) or working in a biotech/health-tech environment.
  • Start-up agility: A proven track record of thriving in fast-paced, ambiguous startup environments in roles requiring high autonomy and ownership.
  • Array and storage formats: Experience with efficient distributed array storage (e.g., xarray, Zarr, TileDB, and TIFF) for both dense and sparse data, and comfort working close to library internals.
  • Workflow orchestration: Experience with orchestration tools such as Dagster, Airflow, Prefect, or database-backed work queues.
  • Frontend and visualization: Experience building or integrating with frontend and visualization tools that make data explorable.
  • Proactive communicator: Ability to translate complex data architecture concepts into clear explanations for scientists, product managers, and engineers alike.

If your strengths lie in just one of these areas and you are passionate about biological data and scalable systems, we highly encourage you to apply!



The Candidate Journey

To be considered, please submit your CV in English.

We believe in a transparent and collaborative interview process. Our goal is to determine whether there is a strong mutual fit. Here is what you can expect after submitting your application:

  • Screening:
  • An initial 30-minute introductory call with our in-house recruiter
  • A 30-minute with the hiring manager to discuss your background, motivations, and the position in more detail.
  • Interviews: Following a successful screening, you will be invited to a series of interviews:
  1. System design (60 min): A discussion of schema design, evolution, and migration; storage formats and performance trade-offs; and API design.
  2. Take-home project: Based on the outcome of the previous call, you will receive a sample system and an accompanying assignment. This assignment covers topics similar to those discussed in the previous system-design interview. You will submit a brief report and Python code. Note: Candidates are welcome to use AI tools during this process, as they would in their daily work at Bioptimus, but they are expected to understand and be able to defend their solutions during the presentation stage.
  3. Assignment presentation and technical Q&A (60 min): You will discuss your take-home assignment with a panel of two or three interviewers. During the first 20 minutes, you will present and demonstrate your solution, followed by a Q&A. You may also briefly introduce your relevant past experience.
  4. Executive Interview (30 min): A discussion with one or more members of our senior leadership team focusing on the company’s vision, cultural fit, shared expectations, and growth potential.

  • Offer: Following the completion of all interviews, our hiring team will make a final decision. Please note that an offer is contingent upon the successful completion of a reference check.
  • Onboarding: We will be happy to welcome you to the team. Once you accept and sign your offer, we will begin the onboarding process.
Why This Is a Unique Opportunity
  • A collaborative and mission-driven work environment at the forefront of biology and AI.
  • The opportunity to help grow our data platform and accelerate discovery in cancer biology.
  • Competitive salary and equity package.
  • Flexible work arrangements, including remote options.



 




We believe that the unique contributions of all Bioptimists create our success. To ensure that our culture continues to incorporate everyone’s perspectives and experience, we never discriminate based on race, religion, national origin, gender identity or expression, sexual orientation, age, or marital, or disability status. Decisions related to hiring are made fairly, and we provide equal employment opportunities to all qualified candidates. We take responsibility for always striving to create an inclusive environment that makes every employee and candidate feel welcome.

Skills Required

  • Deep, production-level Python expertise and commitment to clean, maintainable code
  • Strong hands-on experience with Pydantic, SQLAlchemy, Alembic, object storage abstractions, and FastAPI or similar frameworks
  • Expertise with relational databases such as PostgreSQL, including schema design, indexing, and query optimization
  • Experience designing API surfaces and exposure to programmatic or agent-access protocols such as Model Context Protocol
  • User-centric mindset and commitment to making data infrastructure usable and accessible to non-technical stakeholders
  • Experience handling biological data formats or working in biotech or health-tech
  • Experience thriving in fast-paced, ambiguous startup environments with high autonomy and ownership
  • Experience with xarray, Zarr, TileDB, TIFF, or other efficient distributed array storage formats
  • Experience with Dagster, Airflow, Prefect, or database-backed work queues
  • Experience building or integrating frontend and data visualization tools
  • Ability to explain complex data architecture concepts clearly to scientists, product managers, and engineers
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Paris
33 Employees
Year Founded: 2024

What We Do

We use cutting-edge technology to transform multiscale data into actionable representations to fuel breakthrough discoveries. Join us on our mission to build foundation models that transform biology.

Similar Jobs

GitLab Logo GitLab

Account Executive

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
Remote
Germany
2500 Employees

GitLab Logo GitLab

Account Executive

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
Remote
Germany
2500 Employees

GitLab Logo GitLab

Account Executive

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
Remote
Germany
2500 Employees

GitLab Logo GitLab

Account Executive

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
Remote
Germany
2500 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account