Meld u aan om toegang te krijgen tot alle onderdelen van onze service.
  • Vacatures zoeken
  • Favorieten
  • Maak een cv
    Nieuw
  • Salarissen
  • Inschrijvingen

Senior Site Reliability Engineer (SRE, Compute Node Team)

Vast, Parttime

Nebius

About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The Role We are looking for a Senior Site Reliability Engineer (SRE) to join the Compute Node team at Nebius AI Cloud.

The Compute Node team is responsible for building and operating the cluster scheduler and node-level services that run and manage virtual machines across all cloud regions. This role focuses on Linux systems engineering, virtualization and operational reliability. You will work close to the operating system and hypervisor, shaping how reliability and observability are embedded into the Compute platform. Y our responsibilities will include: - Ensure reliability, availability and performance of compute nodes running VMs - Analyze and debug Linux systems across user space and kernel space, understanding capabilities, limitations and trade-offs at each layer - Troubleshoot complex production issues involving CPU, memory, NUMA, cgroups and scheduling - Work hands-on with virtualization and containerization, primarily using QEMU/KVM and Linux-native technologies - Design and evolve observability as a core capability of the node layer: metrics, logs, traces, alerts, SLIs and SLOs - Lead incident response, root-cause analysis, and postmortems, driving long-term reliability improvements - Collaborate closely with platform, kernel/hypervisor, GPU and infrastructure teams to improve system design and operability We expect you to have: - Strong Linux expertise: - deep understanding of Linux user space and kernel space - knowledge of kernel subsystems (scheduler, memory management, filesystems, cgroups, namespaces) - clear understanding of system boundaries and constraints at different layers - Virtualization experience: - hands-on experience with QEMU/KVM - understanding of VM lifecycle, performance characteristics and failure modes - Containerization knowledge: - practical experience with containers, namespaces and cgroups - strong understanding of resource isolation and control - Strong debugging skills: - ability to reason about complex system failures - structured, hypothesis-driven approach to incident analysis - SRE mindset: - clear understanding of the SRE role in system design and operations - experience building and operating observability stacks, not just consuming them - ability to turn system behavior into actionable reliability signals Nice to Have / Optional: - Experience with Kubernetes internals or node-level components - Hands-on experience with low-level Linux debugging tools (e. g. perf, eBPF, ftrace, strace, kernel crash dumps) - Familiarity with large-scale compute or bare-metal platforms - Contributions to open-source infrastructure or system software - Experience debugging hardware and driver-level issues, including GPUs, NVLink, InfiniBand Benefits & Perks: - Competitive compensation - Career growth and learning opportunities - Flexibility and ownership - Collaborative and innovative culture - Opportunity to work on impactful AI projects - International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer.

Vacature geplaatst op 21 dagen geleden
Soortgelijke banen die u mogelijk interesserenOp basis van de vacature Senior Site Reliability Engineer (SRE, Compute Node Team) in Amsterdam
  •  ...in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration...  ..., we own the hard problems across compute, storage, networking and applied AI....  ...Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers... 
    Voor ouderen
    Senior
    Team
    Fulltime

    Nebius

    Amsterdam
    21 dagen geleden
  •  ...AI/ML infrastructure. Built by engineers, for engineers. From large-...  ...own the hard problems across compute, storage, networking and applied...  ...North America and Israel. Our team of 1,500+ includes hundreds of...  ...and AI R&D. The Role We're an SRE team within DevTools, looking... 
    Voor ouderen
    Senior
    Team
    Fulltime

    Nebius

    Amsterdam
    21 dagen geleden
  •  ...a global trading firm where engineers, traders, and researchers work...  ...The company is seeking a Senior Site Reliability Engineer to support critical...  ...alongside highly skilled technical teams. If you would like to...  ...Qualifications: Strong SRE, Production Engineering, or... 
    Voor ouderen
    Senior
    Team
    Amsterdam
    2 maanden geleden
  •  .../ML infrastructure. Built by engineers, for engineers. From large-scale...  ...own the hard problems across compute, storage, networking and...  ...North America and Israel. Our team of 1,500+ includes hundreds of...  ...multimodal architectures — fast, reliable, and effortless to deploy at... 
    Voor ouderen
    Senior
    Team
    Fulltime

    Nebius

    Amsterdam
    21 dagen geleden
  •  ...They are now seeking a Site Reliability Engineer to join their Eu...  ...years of experience in SRE, DevOps, or Infrastructure...  ...supporting large-scale compute environments. If...  ...networking, and platform teams to optimise resource...  ...standards across thousands of nodes. Diagnose... 
    Team
    Stage
    Op afstand werken
    Per direct beginnen
    Amstelveen
    2 maanden geleden
  •  ...infrastructure. Built by engineers, for engineers. From...  ...hard problems across compute, storage, networking and...  ...and Israel. Our team of 1,500+ includes hundreds...  ...We’re looking for a Site Reliability Engineer to help build...  ...an engineering-first SRE role: you’ll set clear... 
    Team
    Stage
    Fulltime

    Nebius

    Amsterdam
    21 dagen geleden
  •  ...infrastructure. Built by engineers, for engineers. From...  ...the hard problems across compute, storage, networking...  ...America and Israel. Our team of 1,500+ includes hundreds...  ...We’re looking for a Senior HPC Cluster Engineer to...  ...ensuring efficient and reliable operation. We expect you... 
    Voor ouderen
    Senior
    Team
    Fulltime

    Nebius

    Amsterdam
    21 dagen geleden
  •  .../ML infrastructure. Built by engineers, for engineers. From large-scale...  ...own the hard problems across compute, storage, networking and...  ...North America and Israel. Our team of 1,500+ includes hundreds of...  ...automation, operational efficiency. - Reliability & Mission Control —... 
    Voor ouderen
    Senior
    Team
    Met contract
    Fulltime

    Nebius

    Amsterdam
    21 dagen geleden
  •  ...looking for a Software Engineer to build the systems that...  ...Research and Inference team is your customer: today...  .... You will design the engines that manifest the...  ...detect degraded or failed nodes, drain them safely, trigger...  ...automatically. Own reliability of the pipeline:... 
    Voor ouderen
    Senior
    Team
    Stage

    Together AI

    Amsterdam
    4 dagen geleden
  •  ...looking for an experienced SRE with strong background...  ...join our Data Platform team, part of our local data...  ...quant researchers and engineering teams, for all their...  ...scale Support and own reliability of critical services (e...  ...Lake) as well as query engine technologies (e.g.... 
    Team

    IMC

    Amsterdam
    4 dagen geleden
  • € 2.950 - € 4.950 per maand

     ...over efficiënte en duurzame oplossingen. Je groeit door naar senior modelleur of specialiseert je in parametrisch ontwerpen. Over de...  ...om innovatie en maatschappelijke impact. Je komt terecht in een team van doordenkers waar kennisdeling en doorvragen centraal staan.
    Voor ouderen
    Senior
    Team
    Amsterdam
    6 dagen geleden
  •  ...cloud and high performance computing, Atos Group is committed to...  ...space. An L3 support engineer is the most senior engineer in the daily support...  .... The L3 engineering team is responsible for: Supporting...  ...Experience with VSAN ready nodes, VXrails ~ Knowledge of... 
    Voor ouderen
    Senior
    Team
    Fulltime
    Amstelveen
    27 dagen geleden
  •  ...Technical Support Engineer to join our...  ...Enterprise Systems team in Amsterdam....  ...remote and on-site assistance across...  ...closely with senior engineers and teams...  ..., Linux, and SRE to deliver end-...  ...~ BSc in IT/Computer Science preferred...  ...Traders uses reliable market research... 
    Voor ouderen
    Senior
    Team
    Stage
    Op afstand werken

    Flow Traders

    Amsterdam
    3 dagen geleden
  •  ...infrastructure. Built by engineers, for engineers. From...  ...the hard problems across compute, storage, networking...  ...America and Israel. Our team of 1,500+ includes hundreds...  ..., product managers, SRE, infrastructure, networking...  ...operations, and reliability engineering - Previous... 
    Voor ouderen
    Senior
    Team
    Fulltime

    Nebius

    Amsterdam
    20 dagen geleden
  •  ...are seeking a Senior Technology...  ...and execute the reliability and platform strategy...  ...for an SRE team uniquely spanning...  ..., chaos engineering, observability...  ...how whitelabel sites, ad networks,...  ..., and operate reliable services with...  ...Bachelor's degree in Computer Science,... 
    Voor ouderen
    Senior
    Team
    Thuiswerk
    Nachtdienst

    Booking.com

    Amsterdam
    maand geleden
  •  ...we een ervaren DevOps Engineer. PSS is verantwoordelijk...  ...organisatie. Jouw werk als Senior IT Engineering bij ING...  ...binnen een Agile/Scrum team nauw samen met de...  ...ervaring in een DevOps, SRE, Platform Engineering...  ...niveau, bij voorkeur in Computer Science, Cybernetics, Software... 
    Voor ouderen
    Senior
    Team
    Fulltime
    Opdracht
    Thuiswerk
    Amsterdam
    28 dagen geleden
  •  ...choice. At Adyen, everything we do is engineered for ambition.  For our teams, we create an environment with...  ...platforms, of which the biggest is 500+ nodes, with 25k+ CPU cores, 240+ TBs of...  ...flavors (databases, filesystems, compute, etc.) Identify opportunities to... 
    Voor ouderen
    Senior
    Team
    Op afstand werken

    Adyen

    Amsterdam
    4 dagen geleden
  •  ...fastest LLM inference engine with state-of-the-...  .... As a Senior AI Infrastructure Engineer...  ...customers with automated node lifecycle...  ...and non-technical team members ~ Deep experience...  ...high-performance compute, networking, and/or...  ...that scales reliably, fine-tuning and reinforcement... 
    Voor ouderen
    Senior
    Team

    Together AI

    Amsterdam
    4 dagen geleden
  •  ...infrastructure, creating unusual engineering challenges across...  ...systems. As a Senior Full-Stack Software Engineer...  ...Product Engineering team in Amsterdam. This is a...  ...Django REST Framework, Node.js, and Go Building...  ...or master’s degree in Computer Science, Software Engineering... 
    Voor ouderen
    Senior
    Team
    Fulltime
    Thuiswerk

    Surfly

    Amsterdam
    maand geleden
  • € 2.000 per maand

     ...the challenge? As Senior Full-Stack Engineer you will join our Software...  ...& Infrastructure team . You will work...  ...all of it: built in Node.js/TypeScript (Express...  ...corners, and improving reliability. We're a scaleup: we...  ...Bachelor's degree in Computer Science or... 
    Voor ouderen
    Senior
    Team

    Quatt

    Amsterdam
    2 maanden geleden
  •  ...As a DevSecOps Engineer, you'll work closely...  ...closely with Yusuf, our Senior DevOps Engineer,...  ...are secure, reliable, and scalable, while...  ...Tech and Product teams the tools they need...  ...commit), Packer Compute: ECS on Fargate, Lambda...  ...stack: Node.js/TypeScript (NestJS... 
    Voor ouderen
    Senior
    Team
    Stage

    LTD NL B.V.

    Amsterdam
    18 dagen geleden
  • $1,000 per maand

     ...join our Analytics team and develop and...  ...development, software engineering, data handling,...  ...data, establishing nodes and relationships...  ...— Google Earth Engine Models: Implement...  ...collaborating closely with senior research and...  ...field — for example Computer Science,... 
    Voor ouderen
    Senior
    Team
    Stage
    Fulltime
    Op afstand werken
    Thuiswerk

    Laterite

    Amsterdam
    23 dagen geleden
  •  ...increasing demand for large-scale compute, combining cutting-edge...  ...experienced Mechanical Design Engineer looking to shape the mechanical...  ...data centers. You'll work across site evaluation, technical due...  ...construction, and operations teams to support global infrastructure... 
    Voor ouderen
    Senior
    Team
    Stage
    Voor uitvoerders
    Amsterdam
    maand geleden
  • Bedrijfsomschrijving Ben jij een Manager, Senior Manager op het gebied van Identity and...  ...naar een IAM Manager met visie om ons team te versterken en onze klanten te helpen bij...  ...op IT Management, AI, Cybersecurity, Computer Sience etc.)  ~ Woonachtig in Nederland... 
    Voor ouderen
    Senior
    Team

    KPMG Nederland

    Amstelveen
    14 uur geleden
  •  ...built to perform where reliability, resilience and operational...  ...About the Resilience Team The Resilience Team bridges engineering and operations. We identify...  .... The role As a Senior Electrical & Avionics Engineer...  ..., payloads and embedded computers. Develop harnesses,... 
    Voor ouderen
    Senior
    Team
    Fulltime

    DeltaQuad

    Duivendrecht
    maand geleden
  •  ...building next-generation compute environments at scale....  ..., ensuring the reliability, availability and performance...  ...Working closely with senior leadership, construction and engineering teams, you will build operational...  ...standards Own site uptime, reliability, incident... 
    Voor ouderen
    Senior
    Team
    Met contract
    Amsterdam
    2 maanden geleden
  •  ...accelerate the transition to a reliable and renewable energy...  ...the management team. A central part of your...  ...leaders ~A degree in Computer Science, Information...  ...Information Systems, Engineering or a related field, or...  ...Collaborate closely with senior leadership and... 
    Voor ouderen
    Senior
    Team
    Met contract
    Interim

    IT Recruitment

    Amsterdam
    14 dagen geleden
  •  ...looking for an experienced Aerospace Design Engineer to lead the design and development of...  ...in aircraft design, aerodynamics, Computational Fluid Dynamics (CFD), and flight performance...  ...closely with multidisciplinary engineering teams Contribute to the long-term aircraft... 
    Voor ouderen
    Senior
    Team
    Fulltime

    DeltaQuad

    Duivendrecht
    2 maanden geleden
  •  ...You will be part of a growing SAP AI team and work closely with colleagues...  ..., and AI architectures. Advising senior stakeholders on technology strategy,...  ...preferably in information systems, computer science, software engineering, data science, or a related field.... 
    Voor ouderen
    Senior
    Team
    Fulltime
    Thuiswerk

    KPMG Nederland

    Amstelveen
    17 dagen geleden
  •  ...setups , system implementations , engineering development— ensuring they...  ...~ Bachelor's degree in Computer Science, Information Systems,...  ...experience in IT management or senior IT roles, with a track record...  ...stakeholders Lead cross-functional teams including engineers, analysts... 
    Voor ouderen
    Senior
    Team
    Met contract

    Fortaegis Technologies

    Amsterdam
    20 dagen geleden

Wilt u meer vacatures ontvangen?

Abonneer u om vacatures voor Senior Site Reliability Engineer (SRE, Compute Node Team) te ontvangen. Solliciteer als eerste!