AI Engineer (Offensive Security)

Posted Yesterday
Hiring Remotely in London, GBR
In-Office or Remote
Mid level
Cybersecurity
The Role
Build production-quality LLM-powered agentic systems for offensive security: design autonomous hacking agents, backends/APIs/UIs, evaluation harnesses, observability, and gen-AI features; collaborate with consultants and perform applied research.
Summary Generated by Built In

At CovertSwarm, we're doing something few teams get to do: building agentic workflows for real ethical hacking. This is greenfield engineering where you'll work alongside expert security consultants, translate their judgement into LLM-powered systems that are reliable, observable, and backed by rigorous evals.

You'll join a small, new team and help shape not just what we build but how we build it. All this, in a culture of low ego, high trust and constant learning. It's the perfect role for someone excited by real autonomy and technical challenge. 

This is a remote role open to candidates based in EU or Eastern Standard Time (EST) time zones. We gather in person around once per quarter for occasional offsites to kick off new workstreams and align as a team.

We are looking for 3 new AI Engineers to join our growing international team.


What you'll work on

You'll help design, build, and improve AI systems that support real offensive security work. This may include:

  • Autonomous hacking agents that carry out multi-step offensive workflows end to end. You'll design and build the architecture that makes them reliable and reusable.
  • Backend and interfaces that surface these agents to consultants: APIs, services, and UIs that turn raw agent capability into tools consultants can use independently.
  • Evaluation and benchmarking, building realistic target labs and eval harnesses to measure agent performance, and the validation logic that confirms exploits and reduces false positives. All of it built on testable, observable, and maintainable systems.
  • Gen-AI features for consultants, such as chatbot interfaces, text-generation tools, and strategy ideation features that solve real delivery problems.
  • Applied AI research, separating what's genuinely useful from what's hype. You'll have regular research time and an unlimited training budget for continuous learning

The tech stack

Our current stack includes Python, TypeScript, React, LangGraph, Docker and AWS. You don’t need deep experience in all of these, but you should be comfortable building production-quality systems and learning quickly.

Most of the work sits at the intersection of engineering, agentic AI, evaluation, observability, and usable internal tooling.



Who you are

You're a strong engineer who's gone deep on LLMs and agents. You write good quality code, you think rigorously, you take problems and run with them end to end.


What experience you need to be successful

    1. Strong production engineering experience
      You should have proven experience writing high-quality production code and building reliable applications across at least a portion of our tech stack. This includes writing testable and maintainable software, designing systems end to end, working with CI/CD, deploying services, debugging production issues, and making sensible engineering trade-offs.
    2. Experience building and evaluating agentic workflows
      You must have hands-on experience building agentic workflows or LLM-based systems that involve tool use, multi-step reasoning, orchestration, evaluation, or automation.
      We are especially interested in people who have thought deeply about how agents fail, how to measure performance, and how to move beyond impressive demos into reliable systems. This experience may come from professional work, open-source contributions, research, or serious side projects.
    3. Ownership, judgement, and clear communication.
      You’re excited to work on a frontier problem where AI is being applied to real offensive security workflows in new and meaningful ways. You can take a rough idea, shape the approach, make good technical decisions, build and ship the system, and iterate based on feedback. You keep the wider goal in view, test assumptions quickly, adapt as we learn, and communicate clearly throughout, whether you’re documenting an architecture decision or explaining an agent failure mode to a non-AI specialist.

    The Perks 

    Join a team that values both excellence and balance: 

    • True remote flexibility - work from anywhere. 
    • Unlimited training to keep your skills sharp. 
    • Unlimited vacation - because burnout helps no one. 
    • Private medical insurance and pension scheme. 
    • Conference speaking bonuses. 
    • A culture of radical candor, continuous improvement and technical excellence. 

     

    The Culture 

    At CovertSwarm, we take pride in pushing the boundaries of offensive security. Our team consists of passionate and humble professionals who value creativity, technical depth and delivering results that matter. 

    Ready to join the Swarm? 

    Take the next step in your career by applying today. Let’s talk about how your skills, research mindset and offensive capability align with CovertSwarm’s mission to redefine offensive security. 

    Skills Required

    • Strong production engineering experience (writing testable, maintainable code; CI/CD; deploying services; debugging production issues)
    • Hands-on experience building and evaluating agentic workflows or LLM-based systems involving tool use, multi-step reasoning, orchestration, or automation
    • Deep experience with LLMs and agent architectures
    • Comfortable with Python, TypeScript, React, LangGraph, Docker, and AWS
    • Strong ownership, judgement, and clear communication skills
    Am I A Good Fit?
    beta
    Get Personalized Job Insights.
    Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

    The Company
    HQ: London
    60 Employees
    Year Founded: 2020

    What We Do

    YOU DESERVE TO BE HACKED. They say you should keep your friends close, but your enemies closer — and CovertSwarm is your worst nightmare. A specialist red team of ethical hackers and penetration testers, we attack relentlessly from every angle to hit your brand where it hurts again, and again, and again. If there’s a weak spot in your app or website, we’ll find it. If an update exposes a vulnerability, we’ll spot it. And we’ll raise the alarm fast, before anyone can break in.

    Similar Jobs

    Circle Logo Circle

    Solutions Engineer

    Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
    In-Office or Remote
    London, Greater London, England, GBR
    1050 Employees

    NBCUniversal Logo NBCUniversal

    Senior Manager, Sourcing & Procurement

    AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
    Remote or Hybrid
    Bedford, Bedfordshire, England, GBR

    Atlassian Logo Atlassian

    Customer Success Manager

    Cloud • Information Technology • Productivity • Security • Software • App development • Automation
    Remote
    United Kingdom
    11000 Employees

    Atlassian Logo Atlassian

    Marketing Manager

    Cloud • Information Technology • Productivity • Security • Software • App development • Automation
    In-Office or Remote
    London, Greater London, England, GBR
    11000 Employees

    Similar Companies Hiring

    Copia Automation Thumbnail
    Cybersecurity • Industrial
    New York, New York
    50 Employees
    SEON Thumbnail
    Artificial Intelligence • Cybersecurity
    Budapest, Budapest
    415 Employees
    NODA AI Thumbnail
    Artificial Intelligence • Information Technology • Software • Cybersecurity
    Sydney, AU
    54 Employees

    Sign up now Access later

    Create Free Account

    Please log in or sign up to report this job.

    Create Free Account