Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Inference Engineer

£55.6 - £65.8 per hourEstimated

Fuse Energy is a forward-thinking renewable energy startup on a mission to deliver a terawatt of renewable energy - fast. We're combining first-principles thinking with cutting-edge technology to build a radically better energy system. We raised $210M from top-tier investors including Multicoin, Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, Box Group and strategic angels like Nico Rosberg, the Co-Founder of Solana and GPs behind Meta, Revolut, Spotify, Uber and more.

As data centres become one of the largest and fastest-growing sources of electricity demand, Fuse is expanding into high-performance compute infrastructure that sits at the intersection of energy and AI. We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch - and we're looking for the founding engineer to own the latter.

We're looking for a Founding AI Inference Engineer to define and build how Fuse serves AI inference workloads at scale, reporting directly to the CTO. Where our CUDA and GPU engineering hires own kernel-level and hardware performance, this role owns the layer above it: how models actually get served, scaled, and delivered against committed performance targets.

The Opportunity

Fuse is seeing significant demand for data centre capacity across the markets we operate in, primarily for inference. Few companies in the world can pair real power delivery with real compute the way Fuse can, which puts inference serving at the heart of how we turn that advantage into the best offering in the market. That's this role.


Responsibilities

  • Define Fuse's inference serving strategy and architecture from first principles.
  • Design and build the serving stack: request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive inference workloads.
  • Own model-level optimisation strategy for serving - deciding where and how to apply quantisation, distillation, speculative decoding, and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineers.
  • Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents).
  • Translate throughput, latency, and uptime commitments into concrete technical specifications and serving capacity plans.
  • Act as a direct technical owner of inference performance and reliability.
  • Work closely with the CUDA and GPU engineering teams to ensure custom kernels and hardware performance work are integrated cleanly into the serving layer.
  • Set the standards, tooling, and benchmarks this function will run on as it grows.

Requirements

  • 4+ years of experience building or operating large-scale inference serving systems, or equivalent strong project/industry experience.
  • Deep, hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding).
  • Strong systems thinking - able to reason about the full path from incoming request to served response across a large cluster.
  • Comfortable working directly with GPU/CUDA engineers to integrate low-level performance work into a serving system.
  • A track record of making high-stakes architecture calls and owning the outcome.
  • Comfort operating without a playbook - this is a founding role shaping a new function around architecture that's still early-stage, not joining an established one.

Nice to Have

  • Experience with Triton or custom ML inference/training frameworks.
  • Experience with autoscaling or capacity planning for large-scale inference workloads.
  • Exposure to multi-tenant serving or SLA-driven infrastructure.
  • Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system.
  • Familiarity with Kubernetes/Slurm for cluster orchestration.
  • Interest or experience in energy markets, grid systems, or sustainability-focused compute.

Benefits

  • Competitive salary and an equity sign-on bonus.
  • Biannual bonus scheme.
  • Fully expensed tech to match your needs.
  • Breakfast and dinner allowance for office based employees.
Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the AI Inference Engineer in London vacancy
  • £72k - £97k per annumEstimated
     ...We are looking for an AI Inference Engineer to join our growing team. We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL.... 
    Suggested
    Full-time

    perplexity

    London
    1 day ago
  • £58k - £78k per annumEstimated
     ...initiatives through experimentation, causal inference and applied data science. Working at...  ..., collaborating closely with product, engineering, analytics and business teams to deliver...  ...Apply statistical, machine learning and AI techniques to solve complex measurement challenges... 
    Suggested
    Hybrid working

    PlayStation Global

    London
    2 days ago
  •  ...About 9fin 9fin is the AI platform powering global debt markets — the world’s largest...  ...data mining. We are looking for a Senior AI Engineer to accelerate our application of AI...  ...resolution Knowledge Graphs; building and inference, graphRAG Recommendation systems and ranking... 
    Suggested
    Full-time
    Summer work
    Hybrid working
    Flexible hours

    9fin

    London
    3 hours ago
  •  ...Are you an AI engineer who loves taking models out of the lab and putting them into real, high-impact production systems? Do you want to...  ...standalone chatbots Build and maintain data flows required for AI inference, evaluation, and monitoring Optimise AI systems for... 
    Suggested

    esynergy

    London
    more than 2 months ago
  • £58k - £78k per annumEstimated
     ...About 9fin 9fin is the AI platform powering global debt markets — the world’s largest...  ...data mining. We are looking for a Senior AI Engineer to accelerate our application of AI...  ...building feature pipelines, online and batch inference systems, low-latency ranking services, experimentation... 
    Suggested
    Long-term contract
    Full-time
    Temporary
    Summer work
    Hybrid working
    Flexible hours

    9fin

    London
    1 day ago
  • £71k - £95k per annumEstimated
     ...countries. The role You'll build production AI systems that change how our commercial teams work. This is a hands-on engineering role: you'll design, ship and maintain LLM-...  ...fine-tuning, evaluation frameworks or causal inference for measuring intervention impact.... 

    Legatics

    London
    3 days ago
  • £79k - £102k per annumEstimated
     ...Working at CreateFuture CreateFuture is an  AI-native consulting partner  where people do...  ...role and team: A hands-on senior AI engineer embedded in client delivery teams,...  ...implementation of agents, knowledge bases, and model inference in production environments, not limited to... 
    Full-time
    Hybrid working
    On-site
    Remote
    Flexible hours

    CreateFuture

    London
    27 days ago
  • £63k - £84k per annumEstimated
     ...As a Software Engineer on the ML Infrastructure team, you will design and build platforms for...  ...SGLang, TensorRT-LLM, or text-generation-inference. PLEASE NOTE:  Our policy requires a...  ...Scale, our mission is to develop reliable AI systems for the world's most important decisions... 

    Scale AI

    London
    6 days ago
  • £75k per annum

     ...Role : AI Enigneer Location : London Type : Hybrid Salary : Up To £75,000 P/A + Comps The role We're partnering...  ...Requirements Strong experience with Python and SQL, alongside data engineering or software development skills. Hands-on experience with cloud... 
    Permanent
    Hybrid working

    Hunter Bond

    London
    15 days ago
  • $85 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San...  ..., and Jack Dorsey . Position: ML Engineer (Coding Agent Experience) Type: Contract...  ...implementations involving model training , inference systems , MLOps , and LLM applications... 
    Remote job
    Summer work

    Mercor

    London
    a month ago
  • £73k - £106k per annumEstimated
     ...At Graphcore, we’re building the future of AI compute.We’re a team of semiconductor,...  ...intelligence. Job Summary As a research engineer at Graphcore, you will contribute to the...  ...topics, including efficient training and inference, world models, life sciences, reinforcement... 
    Long-term contract
    Visa support
    On-site
    Flexible hours

    Graphcore

    London
    16 days ago
  • £90k - £160k per annum

     ...You have built real AI systems that ship, break, and get fixed under pressure. You...  ...will work at the intersection of reasoning engines, proprietary data, and real market impact....  ...professional investors Build and maintain inference pipelines that are fast, observable, and... 
    Long-term contract
    No agency
    On-site

    Reflexivity

    London
    a month ago
  • £59k - £75k per annumEstimated
     ...software, and applications that power today's AI stack using sustainable technology...  ...Nscale is looking for Senior / Staff AI Engineers to join our core AI team and build the systems...  ...evaluation, and low-latency, high-throughput inference under strict performance and efficiency... 

    Nscale

    London
    a month ago
  • £82k - £108k per annumEstimated
     ...enabling companies to build, train, and serve AI models tailored to their own data,...  ...performance through training and harness engineering. ( blog ) The fine-tuning bottleneck is...  ...product is built on GenAI, architect the inference foundations that capability depends on, and... 

    Fireworks AI

    London
    8 days ago
  • £600 - £750 per day

    Company: INFUSED SOLUTIONS LTD Job Type: Contract, Full Time Salary: £600 - £750/day market rate
    Full-time

    INFUSED SOLUTIONS LTD

    London
    18 days ago
  • £90k - £100k per annum

     ...Job Title: Senior AI and Data Engineer Location: Hybrid Working – London N3 / EC4M Working Hours: Monday to Friday, 35 hour week (Flexitime) Reporting to: Operations & Transformation Director Salary: £90,000 - £100,000 About BKL BKL is a Top 40 accountancy... 
    Hybrid working
    On-site
    Remote
    Monday to Friday
    Flexible hours

    BKL

    London
    a month ago
  •  ...Role: AI Data Engineer Location: 100% remote Start date: May 2026 Length: 12 month contract, with option to extend or convert to full time About: Enterprise manufacturing organisation scaling AI initiatives globally. Looking for a Data Engineer focused on building... 
    Full-time
    Remote
    London
    more than 2 months ago
  • £75k - £95k per annum

     ...Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we...  ...experimentation, training, and production inference, with security, observability, and control...  ...is looking to hire AI Platform Support Engineers to join our EMEA Customer Experience team... 
    Long-term contract
    Hybrid working
    On-site
    Work from home
    Flexible hours
    Shift work
    Weekend work

    Lightning AI

    London
    28 days ago
  • £71k - £95k per annumEstimated
     ...Job Description Help turn cutting edge AI research into real, production grade capabilities used across the firm. You will build...  ...applications that accelerate experimentation while meeting high engineering standards. You will partner closely with AI Researchers and Data... 
    Long-term contract

    JPMorgan Chase & Co.

    London
    3 days ago
  • £60k - £70k per annum

     ...Step into a pivotal role shaping how enterprise organisations design, build and scale AI-ready data platforms. You’ll lead architecture decisions that directly influence cloud transformation, AI adoption and modern digital operating models across complex environments. You... 
    Full-time
    Hybrid working
    Flexible hours

    Cortex IT Recruitment

    Central London
    a month ago
  • £97k - £123k per annumEstimated
     ...We're looking for a Senior AI Engineer / Data Scientist to join our team in London, United Kingdom, in a hybrid working mode. This role is at the core of EPAM’s Data & AI Practice and focuses on building state-of-the-art Generative AI, Agentic AI, and advanced data science... 
    Hybrid working

    EPAM Systems

    London
    29 days ago
  •  ...Contract role: Principal AI Data Engineer Contract Location: London, 5 days onsite weekly Contract Start Date: August 2026 Contract Duration: 4 months Payroll provider:   Rockford Payroll Job Description: Key Responsibilities Develop and evaluate... 
    Long-term contract

    Deloitte - Recruitment

    City of London, Greater London
    1 day ago
  • £57k - £74k per annumEstimated
     ...Description Wavestone is seeking a Data Architect or Data Engineer to support the design and delivery of innovative data solutions...  ...organisations unlock the value of their data by enabling advanced analytics, AI, and data-driven decision-making, while contributing to the... 
    Full-time
    On-site
    Probationary period
    Flexible hours

    Wavestone

    London
    16 days ago
  • £60k - £70k per annum

    Company: ROTHSTEIN RECRUITMENT LTD Job Type: Permanent Salary: £60000.00 - £70000.00
    Permanent

    ROTHSTEIN RECRUITMENT LTD

    London
    16 days ago
  • £80k - £110k per annum

    Company: OPUS RECRUITMENT SOLUTIONS Job Type: Permanent, Full Time Salary: £80000 - £110000/annum
    Permanent
    Full-time

    OPUS RECRUITMENT SOLUTIONS

    London
    10 days ago
  •  ...building a data-driven, operationally efficient platform business. AI will be a structural lever in improving internal productivity,...  ...into how we operate. The Role We are hiring a contract AI Engineer to build, deploy and maintain AI-enabled automations across our core... 

    Alto

    London
    3 days ago
  • £52k - £69k per annumEstimated
     ...Complexio is the intelligence layer for enterprise AI. Our platform builds a connected understanding of how businesses actually operate...  ..., now scaling rapidly across industries. As a Senior AI Engineer with broad expertise, you will be a vital part of our team, developing... 

    Complexio

    London
    3 days ago
  • £56k - £74k per annumEstimated
     ...through safe, impactful and human-centric AI. With more than a decade of experience,...  ...that maximise value. Working closely with engineers, designers, and product managers to ensure...  ...experience. Knowledge in NLP, Bayesian inference, computer vision, deep learning, or causal... 
    Hybrid working
    On-site
    Work from home
    3 days/week

    Faculty

    London
    more than 2 months ago
  •  ...impact within one of our product areas. A core part of the role is designing, running and interpreting experiments, and using causal inference to guide product decisions in complex, real-world settings. You will go beyond simply answering “did this feature work?” and... 
    Long-term contract
    Full-time
    Hybrid working
    On-site
    Flexible hours

    deliveroo

    London
    3 hours ago
  • £60k - £90k per annum

     ...people who are eager to learn and grow, and share our values. Read more about the role and apply. About the role As a Senior AI Engineer at Futurice, you'll be part of our growing AI & Data team, designing and building the intelligent applications and workflows that sit... 
    Permanent
    Visa sponsorship
    On-site
    Remote
    Flexible hours

    Futurice

    London
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Inference Engineer. Be the first to apply!