Senior Full Stack Engineer, Observability

Posted Yesterday
12 Locations
In-Office or Remote
85K-195K Annually
Senior level
Cloud • Software
NetBox Labs makes it easier to build, run, and govern complex networks and infrastructure, for both humans and AI agents
The Role
Build and operate full-stack observability features across Go and Python services, gRPC and REST APIs, data ingestion pipelines, and React/TypeScript dashboards. Design scalable telemetry visualizations, real-time data flows, automated tests, and production reliability practices. Collaborate cross-functionally, participate in on-call, mentor teammates, shape architecture and roadmaps, and use AI-enabled development tools responsibly.
Summary Generated by Built In

NetBox Labs is seeking a Full Stack Engineer to join our rapidly expanding engineering team. We have multiple positions open at different seniority levels across several teams.

About NetBox Labs

NetBox Labs builds the next generation of network automation tools for modern infrastructure teams, with NetBox as the network source of truth at the center. The Observability team builds the products that connect that source of truth to what is actually running on the network:

  • NetBox Discovery finds devices, interfaces, and other network entities.

  • NetBox Assurance brings discovered data into NetBox and helps operators spot drift between intended and actual state, review changes, and decide what to accept.

  • Fleet Management and the Orb agent deploy, configure, and manage the agents that collect discovery and telemetry data from customer networks. These agents use SNMP, gNMI, and other device interfaces.

  • Diode is the ingestion pipeline that moves that data into NetBox reliably.

We're hiring a Full Stack Engineer who can work across the whole path, from the agents and services that collect network data to the interfaces operators use to understand and act on it.

Role overview

You'll join the Observability team and take ownership of features end to end. That covers:

  • Go and Python services and gRPC APIs in the control and data planes.

  • The data flows that carry discovery and telemetry results into NetBox.

  • The React dashboards and interfaces where customers monitor network and device health, explore telemetry, review discovered data, and manage their agent fleet.

You'll work closely with product, design, and other engineering teams, and you'll help run the services the team owns in production.

This role is hands-on. You'll ship features across the stack, improve reliability and performance, and help define the architecture and practices needed to scale network discovery and assurance to large, complex customer environments.

What you'll do

  • Design, build, and operate backend services in Go and Python for discovery, assurance, fleet management, and data ingestion.

  • Define and evolve gRPC and REST APIs with clear, versioned contracts using Protocol Buffers and OpenAPI.

  • Build React and TypeScript monitoring and telemetry dashboards that show device, interface, and network health in real time, with time-series charts, status views, and drill-down from fleet to device to interface.

  • Design dashboard experiences that help operators spot problems fast: sensible defaults, time-range and filter controls, thresholds and alert states, and clear links from a metric to the underlying device in NetBox.

  • Work with backend engineers on the query and aggregation APIs that power dashboards, so they stay fast with large fleets and high-frequency telemetry.

  • Build type-safe integration between frontend and backend using generated API clients, shared schemas, and consistent error and auth handling.

  • Participate in the team's on-call rotation for the services it owns.

  • Add automated tests across the stack (unit, integration, contract, and end-to-end) and enforce quality gates in CI.

  • Collaborate with product managers, designers, customer-facing teams, and other engineering teams. The goal is for features to solve real network operator problems.

  • Use AI-enabled development tools and agentic workflows day to day to speed up design, coding, testing, code review, and incident triage, and help the team adopt them effectively and safely.

  • Review code, mentor teammates, and share best practices for service design, API design, and frontend engineering.

  • Participate in planning processes and help shape the roadmap.

What we're looking for (minimum qualifications)

  • Experience: 5+ years of professional software engineering, with meaningful production experience on both backend and frontend.

  • Backend:

    • Production experience with Go and Python: strong in at least one and working proficiency in the other.

    • Hands-on experience designing and operating gRPC services with Protocol Buffers, including schema evolution and backward compatibility, streaming RPCs, deadlines, interceptors/middleware, and error handling.

    • Experience building distributed, event-driven systems, including message queues (e.g., RabbitMQ or Kafka), asynchronous job processing, and idempotent data ingestion.

  • Frontend (monitoring and telemetry dashboards):

    • Strong React and TypeScript skills, including component composition, state management, and typing best practices.

    • Proven experience building monitoring, observability, or analytics dashboards: time-series charts, heatmaps, status and health views, and drill-down navigation.

    • Hands-on experience with data visualization libraries (e.g., D3, ECharts, Recharts, uPlot, or Visx).

    • Experience rendering large or high-frequency datasets performantly, using techniques such as downsampling, virtualization, and canvas or WebGL rendering.

    • Experience handling real-time data in the browser via WebSockets, server-sent events, or gRPC-Web streaming, along with caching and refresh strategies (e.g., TanStack Query).

    • Modern CSS (TailwindCSS or similar utility-first frameworks), responsive layout, and practical accessibility (WCAG), including color-blind-safe palettes and accessible charts.

    • Automated frontend testing with Jest and React Testing Library, including visual regression testing for charts and dashboards.

  • AI-enabled development:

    • Practical, regular use of AI coding assistants and agents (e.g., Claude Code, Cursor, GitHub Copilot) across the development lifecycle: writing and refactoring code, generating tests, reviewing changes, and writing documentation.

    • Good judgment about when to trust AI output, including verifying generated code, keeping changes reviewable, and protecting sensitive data such as customer network information and credentials.

    • Familiarity with prompting and context techniques that make AI tools effective on a real codebase, such as project instructions, specs, and structured workflows.

  • Ways of working:

    • Good communication skills and a proven ability to work collaboratively in small, cross-functional teams.

    • Comfortable working in a fast-moving environment and contributing to product and technical decisions.

Nice to have

  • Networking protocols and device interfaces:

    • Familiarity with networking protocols and concepts such as TCP/IP, DNS, DHCP, BGP, OSPF, VLANs, and LLDP/CDP.

    • Experience with network management and telemetry interfaces such as SNMP (MIBs, polling, traps), gNMI/gNOI, streaming telemetry.

  • Network automation: Experience with NetBox, network automation frameworks (e.g., NAPALM, Nornir, Netmiko), or building agents that run in customer environments.

  • Telemetry and observability: OpenTelemetry/OTLP, Prometheus, or time-series databases (e.g., Mimir, InfluxDB, ClickHouse), including query languages such as PromQL.

  • Dashboard tooling:

    • Grafana panel or plugin development, or embedding third-party dashboards in a product.

    • User-configurable dashboards (saved views, custom layouts, dashboards as code).

    • Designing alerting and incident UX: thresholds, alert states, annotations, and event timelines.

  • Building with AI:

    • Building AI-powered product features or internal tools using LLM APIs, such as natural-language queries over telemetry, anomaly explanations, or assisted troubleshooting.

    • Experience with the Model Context Protocol (MCP), tool-using agents, or agentic development frameworks and workflows.

    • Evaluating AI features for quality and reliability, such as eval suites and guardrails.

About NetBox Labs:

NetBox Labs helps companies build and manage complex networks. We help customers accelerate network automation by delivering open, composable products and supporting the network automation community.

NetBox Labs is the commercial steward of open source NetBox, the world’s most popular network source of truth, and Orb, the next-generation open source network observability platform. Our products include NetBox Enterprise, a fully supported self-managed NetBox with advanced features, and NetBox Cloud, a secure, scalable, and reliable SaaS edition of NetBox.

NetBox powers thousands of companies, and NetBox Labs is backed by investment from Notable Capital (formerly GGV), Grafana Labs CEO Raj Dutt, Flybridge, IBM, Salesforce Ventures, and Mango Capital.

Our culture and values:
  • We own and solve problems with high attention to detail.

  • Our open source contributors, users, customers & team are all part of our community. When our community wins, we win.

  • We prioritize simplicity and think twice before adding complexity

  • Clear communication helps keep our team aligned and collaborating smoothly.

NetBox Labs is proud to be an equal opportunity employer. We believe diverse teams build better software, and we welcome applicants of every race, color, religion, gender identity, sexual orientation, national origin, age, disability, and veteran status. If you need accommodation at any point in the process, just let us know.

 

Skills Required

  • 5+ years of professional software engineering experience with meaningful production experience across backend and frontend systems
  • Production experience with Go and Python, with strong proficiency in at least one and working proficiency in the other
  • Experience designing and operating gRPC services with Protocol Buffers, including schema evolution, backward compatibility, streaming RPCs, deadlines, middleware, and error handling
  • Experience building distributed, event-driven systems with message queues such as RabbitMQ or Kafka, asynchronous job processing, and idempotent data ingestion
  • Strong React and TypeScript skills, including component composition, state management, and typing best practices
  • Experience building monitoring, observability, or analytics dashboards with time-series charts, heatmaps, health views, and drill-down navigation
  • Hands-on experience with data visualization libraries such as D3, ECharts, Recharts, uPlot, or Visx
  • Experience rendering large or high-frequency datasets using downsampling, virtualization, or canvas/WebGL rendering
  • Experience handling real-time browser data through WebSockets, server-sent events, or gRPC-Web streaming, with caching and refresh strategies
  • Experience with modern CSS, utility-first frameworks such as TailwindCSS, responsive layout, WCAG accessibility, and accessible chart design
  • Automated frontend testing experience with Jest and React Testing Library, including visual regression testing for charts and dashboards
  • Regular practical use of AI coding assistants and agents such as Claude Code, Cursor, or GitHub Copilot across the development lifecycle
  • Ability to verify AI-generated code, keep changes reviewable, protect sensitive data, and use effective prompting and context techniques
  • Strong communication and collaboration skills in small, cross-functional teams
  • Familiarity with networking protocols including TCP/IP, DNS, DHCP, BGP, OSPF, VLANs, and LLDP/CDP
  • Experience with SNMP, gNMI/gNOI, streaming telemetry, or network device interfaces
  • Experience with NetBox, NAPALM, Nornir, Netmiko, or agents running in customer environments
  • Experience with OpenTelemetry, OTLP, Prometheus, time-series databases, or PromQL
  • Grafana panel or plugin development, embedded dashboards, configurable dashboards, or dashboards as code
  • Experience designing alerting and incident user experiences
  • Experience building AI-powered features or internal tools using LLM APIs
  • Experience with Model Context Protocol, tool-using agents, agentic development frameworks, evaluation suites, or guardrails

What the Team is Saying

Natalia Kepets
Chris Veith
Natalie Pastrof
Riley Lewis
Katorey Shinault
Peter Armstrong
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
125 Employees
Year Founded: 2023

What We Do

NetBox Labs simplifies the full infrastructure lifecycle for the world’s most demanding technical environments. As the commercial steward of NetBox, the open-source infrastructure system of record trusted by 10,000+ organizations for more than a decade, NetBox Labs streamlines infrastructure procurement, modeling, deployment and management for both humans and agents. The company’s infrastructure intelligence platform powers business-critical systems at companies like ARM, CoreWeave, J.P. Morgan, Kaiser Permanente, and Riot Games that trust NetBox Labs to manage the networks and infrastructure critical to their business. Headquartered in New York City, NetBox Labs is backed by NGP, Notable Capital, Flybridge, IBM, Salesforce, and Two Sigma. NetBox Cloud, by NetBox Labs, is a fully supported, hosted solution with specific performance and service SLAs and commercial support. NetBox Cloud eliminates the administrative overhead associated with hosting and managing NetBox instances

Why Work With Us

Working at NetBox Labs means stepping on the accelerator. We ship fast, we take ownership, and we swing big. "Impact" is what guides our thoughts and actions. If you want to change the infrastructure industry and add a rocketship to your career, there’s no better place.

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

NetBox Labs Offices

Remote Workspace

Employees work remotely.

Work from almost anywhere — our team is fully remote.

Typical time on-site:
US

Similar Jobs

In-Office or Remote
12 Locations
125 Employees
75K-195K Annually
In-Office or Remote
12 Locations
125 Employees
85K-195K Annually
In-Office or Remote
12 Locations
125 Employees
75K-185K Annually
In-Office or Remote
12 Locations
125 Employees
85K-195K Annually

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account