Staff Engineer, Datacenter Server Lifecycle
Anthropic · San Francisco, CA | New York City · Onsite
Anthropic
Anthropic is an AI safety and research company founded in 2021 by former OpenAI executives, including siblings Dario and Daniela Amodei, who left to pursue a more safety-focused approach to building powerful AI systems. Its flagship product, the Claude family of large language models, competes directly with OpenAI's GPT series and Google's Gemini, and has become widely used in enterprise and developer settings, particularly for coding and reasoning tasks. The company positions itself distinctly within the AI industry by emphasizing interpretability and alignment research alongside commercial deployment, framing itself as a lab racing to build capable frontier models while trying to ensure they remain steerable and beneficial as capabilities scale.
Gig description
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role Anthropic is investing $50 billion in American computing infrastructure, including datacenters custom-built for our workloads, and this role sits at the heart of that effort. As a Staff Engineer on the Datacenter Server Lifecycle team, you will own the end-to-end operational journey of every machine in our datacenters – from initial provisioning and deployment, steady-state operation, maintenance, and repair. This is greenfield work: you will help define the services, tooling, standards and processes that govern how we operate critical hardware at scale, powering frontier model development and serving. You will have the opportunity to define AI-native workflows for datacenter operations and to help drive innovations on efficiency, performance, and reliability. A distinguishing aspect of this role is its deep intersection with security. The machines in our datacenter handle some of the most sensitive workloads in AI – training frontier models and serving millions of users interacting with Claude. Ensuring that every machine in the fleet is trusted, attested, and operating with a verified chain of integrity from the hardware up is a core part of the job, not an afterthought. You will partner closely with our Infrastructure Security team to define and enforce trusted compute standards across the lifecycle, from secure provisioning through end-of-life handling. Key responsibilities Build automation to support datacenter fleets at scale. Define and own the end-to-end system lifecycle strategy – from provisioning and deployment through operation, maintenance, refresh, and decommissioning – and maintain automation and operational procedures for common lifecycle events (e. g. , hardware failures, firmware upgrades, fleet rotations). Partner closely with Infrastructure Security to design and enforce trusted compute standards across the server lifecycle. Work closely with our Networking team to ensure end-to-end connectivity across all sites. Build and maintain tooling to track machine health, configuration, and operational status across the full datacenter fleet. Minimum qualifications Hands-on experience with server hardware, including rack deployment, cabling, troubleshooting, and understanding failure modes at scale. End-to-end understanding of hardware lifecycle management: asset tracking, provisioning workflows, maintenance scheduling, and decommissioning practices. Proficiency in at least one programming language (e. g. , Python, Rust, Go, or Java). Working knowledge of modern cloud infrastructure, including Kubernetes and large scale cloud providers (e. g. AWS, Azure, GCP). Ability to communicate clearly and build consensus with a wide range of stakeholders. Comfort navigating ambiguity and making progress on complex, cross-functional problems. Willingness to travel occasionally to datacenter sites across North America. Preferred qualifications 8+ years of experience in datacenter infrastructure management, or a closely related discipline. Hands-on experience with GPU or AI accelerator hardware (e. g. , NVIDIA A100/H100, Google TPUs, or AWS Trainium) and an understanding of their operational demands. Familiarity with modern provisioning and OS tooling such as LinuxBoot and NixOS Experience building or contributing to datacenter automation or fleet management platforms. Experience building and deploying server operating system distributions across large server fleets. Background in large-scale capacity planning and hardware refresh strategy, ideally at a hyperscaler or large cloud provider. Experience with trusted compute and hardware security concepts such as secure boot, TPM, hardware attestation, and firmware verification — or a strong desire to develop deep expertise in this area. The annual compensation range for this role is listed below. For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role. Annual Salary: $320,000 — $405,000 USD Logistics Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices. Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this. We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team. Your safety matters to us. To protect yourself from potential scams, remember that Anthropic recruiters only contact you from @anthropic. com email addresses. In some cases, we may partner with vetted recruiting agencies who will identify themselves as working on behalf of Anthropic. Be cautious of emails from other domains. Legitimate Anthropic recruiters will never ask for money, fees, or banking information before your first day. If you're ever unsure about a communication, don't click any links—visit anthropic. com/careers directly for confirmed position openings. How we're different We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. And we value impact — advancing our long-term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time. As such, we greatly value communication skills. The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences. Come work with us! Anthropic is a public benefit corporation headquartered in San Francisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues. Guidance on Candidates' AI Usage: Learn about our policy for using AI in our application process.
Requirements to meet
Disclaimer
Suggestions only — review each course yourself to judge whether it meets the role's requirements. Completing a course doesn't guarantee proficiency or that you'll qualify; hiring standards vary by employer.
Skills required
Python (required)
GapCourses that may help you meet this requirement:
PySpark & Python: Hands-On Guide to Data Processing
Coursera · beginner · $49/mo · 5h · ★ 4.4
This course directly covers Python, which appears as a requirement in the Staff Engineer, Datacenter Server Lifecycle posting at a beginner level and can be completed in approximately 5 hours.
Data Analysis with Python
freeCodeCamp · beginner · Free
This beginner course directly builds Python skills, addressing the Python skill gap required for the Staff Engineer, Datacenter Server Lifecycle role.
Pandas
Kaggle Learn · beginner · Free
This course covers Python, which appears as a requirement in the Staff Engineer, Datacenter Server Lifecycle posting at a beginner level.
AWS (required)
GapCourses that may help you meet this requirement:
AWS Fundamentals Specialization
Coursera · beginner · $49/mo · 40h · ★ 4.8
This course covers AWS, which appears as a requirement in the Staff Engineer, Datacenter Server Lifecycle posting at a beginner level and can be completed in approximately 40 hours.
AWS Cloud Technical Essentials
Coursera · beginner · $49/mo · 20h · ★ 4.8
This course covers AWS, which appears as a requirement in the Staff Engineer, Datacenter Server Lifecycle posting at a beginner level and can be completed in approximately 20 hours.
AWS Cloud Solutions Architect Professional Certificate
Coursera · intermediate · $49/mo · 62h · ★ 4.8
This course covers AWS, which appears as a requirement in the Staff Engineer, Datacenter Server Lifecycle posting at an intermediate level and can be completed in approximately 62 hours.
GCP (required)
GapCourses that may help you meet this requirement:
Machine Learning Operations (MLOps) on Google Cloud Specialization
Coursera · intermediate · $49/mo · 16h · ★ 4.0
This course covers GCP, which appears as a requirement in the Staff Engineer, Datacenter Server Lifecycle posting at an intermediate level and can be completed in approximately 16 hours.
Natural Language Processing on Google Cloud
Coursera · advanced · $49/mo · 7h · ★ 4.4
This course covers GCP, which appears as a requirement in the Staff Engineer, Datacenter Server Lifecycle posting at an advanced level and can be completed in approximately 7 hours.
Azure (required)
GapCourses that may help you meet this requirement:
Describe security and compliance concepts
Microsoft Learn · beginner · Free · 48m · ★ 4.8
This course covers Azure, which appears as a requirement in the Staff Engineer, Datacenter Server Lifecycle posting at a beginner level and can be completed in approximately 1 hours.
Embrace responsible AI principles and practices
Microsoft Learn · beginner · Free · 58m · ★ 4.8
This course covers Azure, which appears as a requirement in the Staff Engineer, Datacenter Server Lifecycle posting at a beginner level and can be completed in approximately 1 hours.
Guide AI workload operations with an AI Center of Excellence
Microsoft Learn · beginner · Free · 37m
This course is recommended to address the Azure skill gap for the Staff Engineer, Datacenter Server Lifecycle role, though no specific matched keywords were found to directly link its content to Azure skills.
Kubernetes (required)
GapCourses that may help you meet this requirement:
Kubernetes for the Absolute Beginners with Hands-on Labs
Coursera · beginner · $49 · 10h · ★ 4.7
This course directly covers Kubernetes, which appears as a requirement in the Staff Engineer, Datacenter Server Lifecycle posting at a beginner level and can be completed in approximately 10 hours.
Introduction to project FarmVibes.AI
Microsoft Learn · beginner · Free · 16m
This course covers Kubernetes, which appears as a requirement in the Staff Engineer, Datacenter Server Lifecycle posting at a beginner level and can be completed in approximately 0 hours.
Monitor hybrid virtual machines, containers, and network resources
Microsoft Learn · beginner · Free · 1.5h
This beginner course covers monitoring containers and other hybrid infrastructure resources, providing foundational exposure relevant to the Kubernetes skill needed for the Staff Engineer, Datacenter Server Lifecycle role.
Java (required)
GapCourses that may help you meet this requirement:
Object Oriented Programming in Java Specialization
Coursera · beginner · $49/mo · 115h · ★ 4.6
This course directly covers Java, which appears as a requirement in the Staff Engineer, Datacenter Server Lifecycle posting at a beginner level and can be completed in approximately 115 hours.
Java Programming for Beginners
Coursera · beginner · $49/mo · 12h · ★ 4.6
This course directly covers Java, which appears as a requirement in the Staff Engineer, Datacenter Server Lifecycle posting at a beginner level and can be completed in approximately 12 hours.
Monitoring Java applications on Azure
Microsoft Learn · intermediate · Free · 1h · ★ 4.8
This intermediate course directly addresses the Java skill gap for the Staff Engineer, Datacenter Server Lifecycle role by focusing on Java, as confirmed by the matched keyword 'java'.
Rust (required)
GapData Center Operations (required)
GapCourse that may help you meet this requirement:
Datacenter integration for Azure Stack Hub
Microsoft Learn · advanced · Free · 1h · ★ 4.8
This course covers Data Center Operations, which appears as a requirement in the Staff Engineer, Datacenter Server Lifecycle posting at an advanced level and can be completed in approximately 1 hours.