Microarchitect / RTL Design - Memory Subsystem

Posted 2 Days Ago
Be an Early Applicant
2 Locations
In-Office
Senior level
Artificial Intelligence • Hardware • Information Technology • Machine Learning
The Role
Own the microarchitecture and synthesizable RTL design of production memory subsystems for AI ASICs. Design DDR, LPDDR, and HBM controllers, cache hierarchies, memory-side accelerators, scheduling, ECC, QoS, and interfaces. Drive RTL through lint, CDC, synthesis, timing, area, and power closure; collaborate with verification, software, ML, and physical design teams. Support FPGA prototyping and AI-driven design automation using SystemVerilog and Python.
Summary Generated by Built In
About Architect

Architect is a frontier AI lab for chip design. We build AI models and tools for on-demand custom ASICs at scale. Our goal is to co-design custom ASICs alongside evolving ML workloads, and enable a new era of domain-specific chips that unlock capabilities impossible with current hardware paradigms. Born out of Stanford Research, our team blends AI with Silicon with a founding team from Anthropic, Google DeepMind, Meta SuperIntelligence, xAI, Apple and Intel.

What You’ll Do

As a Founding Member of the Technical Staff on the RTL Design team at Architect, you’ll own the AI-driven microarchitecture and RTL design of the memory subsystem going into production silicon. You will define, drive, and revise the block-level micro-architecture specification for memory controllers, memory hierarchy management, and memory-side accelerators — ensuring maximum bandwidth utilization, minimal latency, and efficient power delivery for compute-intensive ML workloads.

Core Responsibilities
  • Own the memory subsystem RTL end-to-end: from DDR/HBM controller design through code generation, lint, CDC, synthesis, and timing closure using our AI-driven design flow.

  • Design and implement memory controllers: including DDR5/LPDDR5X PHY-side controller logic, HBM3/HBM3E pseudo-channel controllers, command scheduling (open-page/close-page policies, bank-level parallelism), refresh management, and ECC/RAS engines.

  • Architect the memory hierarchy: including multi-level cache controllers, scratchpad memory managers, coherency protocol engines (where applicable), prefetch engines, and bandwidth partitioning/QoS mechanisms to serve diverse traffic profiles from ML accelerator datapaths.

  • Design memory-side accelerators: near-memory compute logic, scatter-gather DMA engines, address translation/remapping units, compression/decompression engines co-located with memory interfaces, and intelligent prefetchers tuned for ML access patterns.

  • Work directly with the principal architect to refine microarchitectural specs, resolve implementation trade-offs (bandwidth vs. latency vs. area vs. power), and feed area/timing/power realities back into the architecture and internal AI systems.

  • Define and maintain interface specifications: DDR PHY interfaces (DFI), HBM PHY interfaces, on-chip SRAM interfaces, AXI/ACE/CHI for memory-facing fabric ports, and custom interfaces for near-memory accelerator datapaths.

  • Build and maintain RTL infrastructure for our in-house AI-driven flow: design automation scripts, regression flows, lint/CDC waivers, and integration collateral for the memory subsystem.

  • Close collaboration with DV: Support verification bring-up with memory timing models, protocol-compliant BFMs, SVA assertions for JEDEC protocol compliance, coverage plans targeting worst-case scheduling scenarios, and architectural documentation for verification closure.

  • Close collaboration with SW and ML: Support and guide our SW and ML experts to revise and improve our in-house AI flow based on your memory subsystem domain expertise — particularly around workload-driven memory access pattern optimization.

  • Support FPGA prototyping on Xilinx for early functional validation of memory controllers, including bring-up with DDR MIG IPs and HBM validation platforms.

What We’d Like to SeeRequired Qualifications
  • Degree: Bachelor’s, Master’s, or PhD in Electrical Engineering, Computer Engineering, or a closely related field.

  • Experience: 5+ years (10+ preferred) in RTL design with at least one advanced-node tapeout experience involving memory subsystems (DDR/LPDDR/HBM controllers, cache hierarchies, or memory-intensive SoC subsystems).

  • Memory Interface Expertise: Deep familiarity with JEDEC memory standards — DDR5/LPDDR5X command/address protocols, timing parameters, training sequences, and/or HBM2E/HBM3 pseudo-channel architecture, stack addressing, and interleaving schemes.

  • Memory Controller Design: Hands-on experience designing or owning memory controller blocks including command schedulers, bank state machines, refresh engines (per-bank, fine-granularity), read/write turnaround optimization, and PHY interface timing (DFI or proprietary).

  • Memory Hierarchy Architecture: Experience with multi-level cache design (tag/data arrays, replacement policies, coherence protocols), scratchpad controllers, or unified memory architectures with partitioning and QoS.

  • SystemVerilog: Clear, synthesizable, lint-clean RTL with strong design habits — parameterization for multi-standard support (DDR5/HBM3), modularity for channel/pseudo-channel instantiation, and configurability for different capacity/bandwidth targets.

  • Block-Level Depth: Hands-on experience with SRAM controllers and arbiters, bank conflict resolution, address hashing/interleaving, ECC encode/decode engines, and high-bandwidth data movement between on-chip and off-chip memory.

  • SoC Methodology: Solid grasp of synthesis, timing constraints, clock domain crossings (PHY-to-controller domain, multi-frequency memory interfaces), reset strategies, AMBA protocols (AXI, ACE, CHI), and power management for memory subsystems.

  • Python: Strong skills for design automation, performance modeling, regression infrastructure, and tooling.

  • PPA Ownership: Experience taking a memory controller or cache subsystem from RTL through synthesis and working with PD teams on timing/area/power closure — particularly for high-frequency controller logic and wide data buses.

Bonus Qualifications
  • Experience with HBM integration: interposer-level considerations, PHY calibration, thermal management impacts on refresh.

  • Familiarity with CXL memory pooling, Type 3 device controllers, or disaggregated memory architectures.

  • Near-memory or processing-in-memory (PIM) design experience.

  • Low-power design techniques: DVFS-aware memory scheduling, partial-array self-refresh, clock gating of idle channels, power gating of unused banks.

  • FPGA prototyping experience (Xilinx Vivado/Vitis) with DDR MIG or HBM subsystem IP integration.

  • SVA assertions for JEDEC protocol compliance (command sequencing, timing parameter checking, training state machines).

  • Prior IP building and delivery experience for DDR/LPDDR controllers, HBM controllers, or cache subsystem IPs.

  • Performance modeling: experience building or using cycle-accurate memory system simulators (e.g., DRAMSim, Ramulator) to validate microarchitectural decisions.

  • Domain-specific research contributions: publications or patents in memory systems, memory scheduling algorithms, or memory-centric compute architectures for ML workloads.

Why Architect

You’ll join a founding team building the future of chip design at the intersection of AI and silicon. Your memory subsystem expertise will directly shape production ASICs — enabling the bandwidth and efficiency that ML workloads demand — and influence how AI transforms hardware development from spec to tapeout.

Skills Required

  • Bachelor's, Master's, or PhD in Electrical Engineering, Computer Engineering, or a closely related field
  • 5+ years of RTL design experience
  • At least one advanced-node tapeout involving memory subsystems
  • Deep familiarity with DDR5, LPDDR5X, HBM2E, or HBM3 memory standards and architectures
  • Hands-on experience designing or owning memory controller blocks
  • Experience with multi-level caches, scratchpad controllers, or unified memory architectures
  • Strong synthesizable SystemVerilog RTL design skills
  • Experience with SRAM controllers, arbiters, bank conflict resolution, address interleaving, ECC, and high-bandwidth memory data movement
  • Understanding of synthesis, timing constraints, CDC, reset strategies, AMBA protocols, and memory power management
  • Strong Python skills for design automation, performance modeling, regression infrastructure, and tooling
  • Experience taking memory controllers or cache subsystems through synthesis and PPA closure
  • HBM integration experience, including interposer considerations, PHY calibration, or thermal impacts on refresh
  • Familiarity with CXL memory pooling, Type 3 device controllers, or disaggregated memory architectures
  • Near-memory or processing-in-memory design experience
  • Low-power memory design techniques such as DVFS-aware scheduling, clock gating, or power gating
  • Xilinx Vivado/Vitis FPGA prototyping with DDR MIG or HBM subsystem IP
  • SVA assertions for JEDEC protocol compliance
  • Prior IP delivery experience for DDR, LPDDR, HBM, or cache subsystem controllers
  • Experience with cycle-accurate memory simulators such as DRAMSim or Ramulator
  • Publications or patents in memory systems, scheduling algorithms, or memory-centric ML architectures
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Year Founded: 2023

What We Do

an AI lab for compute infrastructure, starting with chip design

Similar Jobs

Sprout Social Logo Sprout Social

Customer Success Manager

Marketing Tech • Social Media • Software • Analytics • Business Intelligence
Easy Apply
Remote or Hybrid
US
1400 Employees
92K-153K Annually
Hybrid
2 Locations
289097 Employees
Hybrid
8 Locations
289097 Employees

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Sales Support Associate III

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
Ontario, CA, USA
16000 Employees
15-22 Hourly

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account