Senior Site Reliability Engineer (SRE, Compute Node Team)
Nebius
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D. The Role We are looking for a Senior Site Reliability Engineer (SRE) to join the Compute Node team at Nebius AI Cloud.
The Compute Node team is responsible for building and operating the cluster scheduler and node-level services that run and manage virtual machines across all cloud regions. This role focuses on Linux systems engineering, virtualization and operational reliability. You will work close to the operating system and hypervisor, shaping how reliability and observability are embedded into the Compute platform. Y our responsibilities will include: - Ensure reliability, availability and performance of compute nodes running VMs - Analyze and debug Linux systems across user space and kernel space, understanding capabilities, limitations and trade-offs at each layer - Troubleshoot complex production issues involving CPU, memory, NUMA, cgroups and scheduling - Work hands-on with virtualization and containerization, primarily using QEMU/KVM and Linux-native technologies - Design and evolve observability as a core capability of the node layer: metrics, logs, traces, alerts, SLIs and SLOs - Lead incident response, root-cause analysis, and postmortems, driving long-term reliability improvements - Collaborate closely with platform, kernel/hypervisor, GPU and infrastructure teams to improve system design and operability We expect you to have: - Strong Linux expertise: - deep understanding of Linux user space and kernel space - knowledge of kernel subsystems (scheduler, memory management, filesystems, cgroups, namespaces) - clear understanding of system boundaries and constraints at different layers - Virtualization experience: - hands-on experience with QEMU/KVM - understanding of VM lifecycle, performance characteristics and failure modes - Containerization knowledge: - practical experience with containers, namespaces and cgroups - strong understanding of resource isolation and control - Strong debugging skills: - ability to reason about complex system failures - structured, hypothesis-driven approach to incident analysis - SRE mindset: - clear understanding of the SRE role in system design and operations - experience building and operating observability stacks, not just consuming them - ability to turn system behavior into actionable reliability signals Nice to Have / Optional: - Experience with Kubernetes internals or node-level components - Hands-on experience with low-level Linux debugging tools (e. g. perf, eBPF, ftrace, strace, kernel crash dumps) - Familiarity with large-scale compute or bare-metal platforms - Contributions to open-source infrastructure or system software - Experience debugging hardware and driver-level issues, including GPUs, NVLink, InfiniBand Benefits & Perks: - Competitive compensation - Career growth and learning opportunities - Flexibility and ownership - Collaborative and innovative culture - Opportunity to work on impactful AI projects - International environment and talented teams What's it like to work at Nebius: Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI Equal Opportunity Statement: Nebius is an equal opportunity employer.
- ...in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration... ..., we own the hard problems across compute, storage, networking and applied AI.... ...Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers...Voor ouderenSeniorTeamFulltime
- ...AI/ML infrastructure. Built by engineers, for engineers. From large-... ...own the hard problems across compute, storage, networking and applied... ...North America and Israel. Our team of 1,500+ includes hundreds of... ...and AI R&D. The Role We're an SRE team within DevTools, looking...Voor ouderenSeniorTeamFulltime
- ...a global trading firm where engineers, traders, and researchers work... ...The company is seeking a Senior Site Reliability Engineer to support critical... ...alongside highly skilled technical teams. If you would like to... ...Qualifications: Strong SRE, Production Engineering, or...Voor ouderenSeniorTeam
- .../ML infrastructure. Built by engineers, for engineers. From large-scale... ...own the hard problems across compute, storage, networking and... ...North America and Israel. Our team of 1,500+ includes hundreds of... ...multimodal architectures — fast, reliable, and effortless to deploy at...Voor ouderenSeniorTeamFulltime
- ...They are now seeking a Site Reliability Engineer to join their Eu... ...years of experience in SRE, DevOps, or Infrastructure... ...supporting large-scale compute environments. If... ...networking, and platform teams to optimise resource... ...standards across thousands of nodes. Diagnose...TeamStageOp afstand werkenPer direct beginnen
- ...infrastructure. Built by engineers, for engineers. From... ...hard problems across compute, storage, networking and... ...and Israel. Our team of 1,500+ includes hundreds... ...We’re looking for a Site Reliability Engineer to help build... ...an engineering-first SRE role: you’ll set clear...TeamStageFulltime
- ...infrastructure. Built by engineers, for engineers. From... ...the hard problems across compute, storage, networking... ...America and Israel. Our team of 1,500+ includes hundreds... ...We’re looking for a Senior HPC Cluster Engineer to... ...ensuring efficient and reliable operation. We expect you...Voor ouderenSeniorTeamFulltime
- .../ML infrastructure. Built by engineers, for engineers. From large-scale... ...own the hard problems across compute, storage, networking and... ...North America and Israel. Our team of 1,500+ includes hundreds of... ...automation, operational efficiency. - Reliability & Mission Control —...Voor ouderenSeniorTeamMet contractFulltime
- ...looking for a Software Engineer to build the systems that... ...Research and Inference team is your customer: today... .... You will design the engines that manifest the... ...detect degraded or failed nodes, drain them safely, trigger... ...automatically. Own reliability of the pipeline:...Voor ouderenSeniorTeamStage
- ...looking for an experienced SRE with strong background... ...join our Data Platform team, part of our local data... ...quant researchers and engineering teams, for all their... ...scale Support and own reliability of critical services (e... ...Lake) as well as query engine technologies (e.g....Team
€ 2.950 - € 4.950 per maand
...over efficiënte en duurzame oplossingen. Je groeit door naar senior modelleur of specialiseert je in parametrisch ontwerpen. Over de... ...om innovatie en maatschappelijke impact. Je komt terecht in een team van doordenkers waar kennisdeling en doorvragen centraal staan.Voor ouderenSeniorTeam- ...cloud and high performance computing, Atos Group is committed to... ...space. An L3 support engineer is the most senior engineer in the daily support... .... The L3 engineering team is responsible for: Supporting... ...Experience with VSAN ready nodes, VXrails ~ Knowledge of...Voor ouderenSeniorTeamFulltime
- ...Technical Support Engineer to join our... ...Enterprise Systems team in Amsterdam.... ...remote and on-site assistance across... ...closely with senior engineers and teams... ..., Linux, and SRE to deliver end-... ...~ BSc in IT/Computer Science preferred... ...Traders uses reliable market research...Voor ouderenSeniorTeamStageOp afstand werken
- ...infrastructure. Built by engineers, for engineers. From... ...the hard problems across compute, storage, networking... ...America and Israel. Our team of 1,500+ includes hundreds... ..., product managers, SRE, infrastructure, networking... ...operations, and reliability engineering - Previous...Voor ouderenSeniorTeamFulltime
- ...are seeking a Senior Technology... ...and execute the reliability and platform strategy... ...for an SRE team uniquely spanning... ..., chaos engineering, observability... ...how whitelabel sites, ad networks,... ..., and operate reliable services with... ...Bachelor's degree in Computer Science,...Voor ouderenSeniorTeamThuiswerkNachtdienst
- ...we een ervaren DevOps Engineer. PSS is verantwoordelijk... ...organisatie. Jouw werk als Senior IT Engineering bij ING... ...binnen een Agile/Scrum team nauw samen met de... ...ervaring in een DevOps, SRE, Platform Engineering... ...niveau, bij voorkeur in Computer Science, Cybernetics, Software...Voor ouderenSeniorTeamFulltimeOpdrachtThuiswerk
- ...choice. At Adyen, everything we do is engineered for ambition. For our teams, we create an environment with... ...platforms, of which the biggest is 500+ nodes, with 25k+ CPU cores, 240+ TBs of... ...flavors (databases, filesystems, compute, etc.) Identify opportunities to...Voor ouderenSeniorTeamOp afstand werken
- ...fastest LLM inference engine with state-of-the-... .... As a Senior AI Infrastructure Engineer... ...customers with automated node lifecycle... ...and non-technical team members ~ Deep experience... ...high-performance compute, networking, and/or... ...that scales reliably, fine-tuning and reinforcement...Voor ouderenSeniorTeam
- ...infrastructure, creating unusual engineering challenges across... ...systems. As a Senior Full-Stack Software Engineer... ...Product Engineering team in Amsterdam. This is a... ...Django REST Framework, Node.js, and Go Building... ...or master’s degree in Computer Science, Software Engineering...Voor ouderenSeniorTeamFulltimeThuiswerk
€ 2.000 per maand
...the challenge? As Senior Full-Stack Engineer you will join our Software... ...& Infrastructure team . You will work... ...all of it: built in Node.js/TypeScript (Express... ...corners, and improving reliability. We're a scaleup: we... ...Bachelor's degree in Computer Science or...Voor ouderenSeniorTeam- ...As a DevSecOps Engineer, you'll work closely... ...closely with Yusuf, our Senior DevOps Engineer,... ...are secure, reliable, and scalable, while... ...Tech and Product teams the tools they need... ...commit), Packer Compute: ECS on Fargate, Lambda... ...stack: Node.js/TypeScript (NestJS...Voor ouderenSeniorTeamStage
$1,000 per maand
...join our Analytics team and develop and... ...development, software engineering, data handling,... ...data, establishing nodes and relationships... ...— Google Earth Engine Models: Implement... ...collaborating closely with senior research and... ...field — for example Computer Science,...Voor ouderenSeniorTeamStageFulltimeOp afstand werkenThuiswerk- ...increasing demand for large-scale compute, combining cutting-edge... ...experienced Mechanical Design Engineer looking to shape the mechanical... ...data centers. You'll work across site evaluation, technical due... ...construction, and operations teams to support global infrastructure...Voor ouderenSeniorTeamStageVoor uitvoerders
- Bedrijfsomschrijving Ben jij een Manager, Senior Manager op het gebied van Identity and... ...naar een IAM Manager met visie om ons team te versterken en onze klanten te helpen bij... ...op IT Management, AI, Cybersecurity, Computer Sience etc.) ~ Woonachtig in Nederland...Voor ouderenSeniorTeam
- ...built to perform where reliability, resilience and operational... ...About the Resilience Team The Resilience Team bridges engineering and operations. We identify... .... The role As a Senior Electrical & Avionics Engineer... ..., payloads and embedded computers. Develop harnesses,...Voor ouderenSeniorTeamFulltime
- ...building next-generation compute environments at scale.... ..., ensuring the reliability, availability and performance... ...Working closely with senior leadership, construction and engineering teams, you will build operational... ...standards Own site uptime, reliability, incident...Voor ouderenSeniorTeamMet contract
- ...accelerate the transition to a reliable and renewable energy... ...the management team. A central part of your... ...leaders ~A degree in Computer Science, Information... ...Information Systems, Engineering or a related field, or... ...Collaborate closely with senior leadership and...Voor ouderenSeniorTeamMet contractInterim
- ...looking for an experienced Aerospace Design Engineer to lead the design and development of... ...in aircraft design, aerodynamics, Computational Fluid Dynamics (CFD), and flight performance... ...closely with multidisciplinary engineering teams Contribute to the long-term aircraft...Voor ouderenSeniorTeamFulltime
- ...You will be part of a growing SAP AI team and work closely with colleagues... ..., and AI architectures. Advising senior stakeholders on technology strategy,... ...preferably in information systems, computer science, software engineering, data science, or a related field....Voor ouderenSeniorTeamFulltimeThuiswerk
- ...setups , system implementations , engineering development— ensuring they... ...~ Bachelor's degree in Computer Science, Information Systems,... ...experience in IT management or senior IT roles, with a track record... ...stakeholders Lead cross-functional teams including engineers, analysts...Voor ouderenSeniorTeamMet contract
Wilt u meer vacatures ontvangen?
Abonneer u om vacatures voor Senior Site Reliability Engineer (SRE, Compute Node Team) te ontvangen. Solliciteer als eerste!
- senior mechanical engineer Amsterdam
- senior design engineer Amsterdam
- senior automation engineer Amsterdam
- senior ict engineer Amsterdam
- salaris senior engineer Amsterdam
- senior validation engineer Amsterdam
- senior data engineer Amsterdam
- senior cloud engineer Amsterdam
- senior system engineer Amsterdam
- senior security engineer Amsterdam
