Sciforium

HQ
San Francisco
Total Offices: 2
7 Total Employees
Year Founded: 2024

Jobs at Sciforium

Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.

Recently posted jobs

2 Days AgoSaved
Hybrid
San Francisco, CA, USA
Artificial Intelligence • Information Technology • Software
The GPU Kernel Engineer will design and optimize GPU kernels for large-scale AI systems, integrating them into ML frameworks and collaborating across teams to enhance performance.
29 Days AgoSaved
In-Office
San Francisco, CA, USA
Artificial Intelligence • Information Technology • Software
Develop and pre-train large byte-native and multimodal foundation models. Implement architectures, training objectives, optimization methods, and stable scaling recipes. Build production-grade training infrastructure, run ablations, analyze training dynamics, and improve model quality, efficiency, reliability, and scalability across GPU-based distributed environments.
One Month AgoSaved
In-Office
San Francisco, CA, USA
Artificial Intelligence • Information Technology • Software
Develop the core systems for multimodal AI models, focusing on building model serving platforms and enhancing developer experience with tools and APIs.
One Month AgoSaved
In-Office
San Francisco, CA, USA
Artificial Intelligence • Information Technology • Software
Develop the model serving platform for multimodal AI models, optimizing performance, collaborating with researchers, and ensuring reliable production infrastructure.
One Month AgoSaved
In-Office
San Francisco, CA, USA
Artificial Intelligence • Information Technology • Software
Lead advanced AI/ML research initiatives, developing novel algorithms and optimizing large-scale training systems. Mentor junior researchers and produce impactful research outputs.
One Month AgoSaved
In-Office
San Francisco, CA, USA
Artificial Intelligence • Information Technology • Software
The role involves optimizing and maintaining the software stack for large-scale AI training, focusing on CUDA/ROCm, JAX, and PyTorch to ensure efficiency and stability in distributed training systems.
One Month AgoSaved
In-Office
San Francisco, CA, USA
Artificial Intelligence • Information Technology • Software
Design, deploy, and operate the complete networking stack for large GPU clusters: RDMA fabrics (InfiniBand/RoCE), frontend/storage networks, security, cross-site/cloud connectivity, automation, telemetry, and performance debugging to ensure the fabric never bottlenecks training and inference workloads.
One Month AgoSaved
In-Office
2 Locations
Artificial Intelligence • Information Technology • Software
Own physical health and infrastructure of GPU clusters: respond to hardware outages, monitor GPU thermals/power, coordinate vendor RMAs and data center work, rack and bring up nodes, maintain Linux OS across bare-metal servers, secure networking and SSH, and manage directory and network storage services to keep compute environment stable for research and product teams.
One Month AgoSaved
In-Office
San Francisco, CA, USA
Artificial Intelligence • Information Technology • Software
Owner of the GPU cluster software stack: build/version OS images, drivers, and provisioning pipelines; automate validation, upgrades, and self-healing; operate GPU-enabled Kubernetes and Slurm; maintain driver/runtime compatibility and tuned PyTorch/JAX environments; debug deep accelerator and networking issues and run observability/benchmarking tooling.