Storage Rack Infrastructure Automation & Cluster Bring-Up - Hive Program

Apply Now ↗

About this role

Company Description

Sandisk understands how people and businesses consume data and we relentlessly innovate to deliver solutions that enable today’s needs and tomorrow’s next big ideas. With a rich history of groundbreaking innovations in Flash and advanced memory technologies, our solutions have become the beating heart of the digital world we’re living in and that we have the power to shape.

Sandisk meets people and businesses at the intersection of their aspirations and the moment, enabling them to keep moving and pushing possibility forward. We do this through the balance of our powerhouse manufacturing capabilities and our industry-leading portfolio of products that are recognized globally for innovation, performance and quality.

Sandisk has two facilities recognized by the World Economic Forum as part of the Global Lighthouse Network for advanced 4IR innovations. These facilities were also recognized as Sustainability Lighthouses for breakthroughs in efficient operations. With our global reach, we ensure the global supply chain has access to the Flash memory it needs to keep our world moving forward.

Job Description

Hive is a swarm of hundreds of identical storage nodes (Marvell/XSight DPU + SSD), each running a full Ceph plane: OSD, Monitor (MON), Manager (MGR), MDS, plus SeaStore and the NFS front end. Standing up, re-configuring, and recovering a cluster of this size by hand does not scale. We need an engineer who owns the automated hardware discovery and Ceph role-assignment pipeline: the flow that inventories every node and its hardware, decides which daemons each node should run, and drives the cluster from bare metal to a serving state. The flow must run fully autonomously by default, with a clean manual-override path for lab, bring-up, and failure-injection scenarios.

This is a specialized automation and lab-hardware-discovery discipline that the Hive team does not currently have dedicated ownership for. It is directly on the critical path for every test cluster, every silicon bring-up, and every customer-shaped deployment.

What you'll own

  • Automated hardware discovery. Detect nodes as they power on and enumerate their hardware (DPU/SoC model, SSD media, DRAM, RNIC/network ports, BMC) using out-of-band and in-band inventory (BMC/Redfish/IPMI, PXE/DHCP boot, cloud-init/first-boot agents). Produce a single authoritative machine inventory that the rest of the pipeline consumes.
  • Role assignment and cluster composition. Given the discovered inventory, decide and apply which Ceph roles each node runs — OSD, MON, MGR, MDS (and the NFS gateway) — encoded as declarative placement specifications. Own MON quorum sizing and placement, MGR redundancy, MDS/gateway placement, and CRUSH map / failure-domain layout so data and metadata land correctly across the swarm.
  • Bare-metal → serving bring-up. Drive the end-to-end sequence: node provisioning, OS/image deploy, cluster bootstrap, daemon deployment via the Ceph orchestrator (cephadm-style), ceph-volume-style OSD provisioning on the SSD media, and health convergence to HEALTH_OK.
  • Autonomous and manual modes. Make the default path zero-touch (a rack powers on and self-assembles into a healthy cluster), while exposing deterministic manual controls to pin roles, hold a node out, force a specific topology, or reproduce a customer/lab configuration for testing.
  • Lifecycle and recovery automation. Node add/remove, drain and rebalance, daemon replacement, MON re-quorum after loss, MDS/OSD failover validation, and re-discovery after re-imaging — integrated with Hive's Fast Recovery + BMC work.
  • Reconciliation and drift control. Continuously compare declared desired state against observed cluster state and converge — the same idea Ceph's orchestrator applies to service specs — with clear reporting when reality diverges from intent.
  • CI/lab integration. Wire the pipeline into automated test so any commit can spin up a correctly-composed multi-node cluster on real hardware and tear it down cleanly.

Qualifications

Core (must-have):

  • Deep experience in lab hardware discovery, inventory, and automated provisioning at fleet scale.
  • Bare-metal automation: PXE/DHCP boot, BMC out-of-band management (Redfish/IPMI), image/OS deployment, cloud-init/first-boot.
  • Infrastructure-as-code and config automation (e.g., Ansible/Terraform-class tooling) with a declarative, reconcile-to-desired-state mindset.
  • Strong scripting/automation (Python and shell) and CI systems (Jenkins/GitLab CI or equivalent).
  • Comfort designing systems that are autonomous by default with reliable manual override.

Strongly preferred (or ramp-up expected):

  • Ceph operational knowledge: OSD/MON/MGR/MDS roles, the orchestrator/cephadm model, placement specs, ceph-volume, CRUSH maps and failure domains, MON quorum, and cluster health/lifecycle.
  • Distributed-systems fluency: quorum/Paxos intuition, rebalance/recovery behavior, failure-domain reasoning.
  • Networking bring-up for storage fabrics (RoCEv2/Ethernet, port discovery).
  • Familiarity with DPU/SoC-based nodes and constrained-node environments.

Additional Information

Sandisk thrives on the power and potential of diversity. As a global company, we believe the most effective way to embrace the diversity of our customers and communities is to mirror it from within. We believe the fusion of various perspectives results in the best outcomes for our employees, our company, our customers, and the world around us. We are committed to an inclusive environment where every individual can thrive through a sense of belonging, respect and contribution.

Sandisk is committed to offering opportunities to applicants with disabilities and ensuring all candidates can successfully navigate our careers website and our hiring process. Please contact us at [email protected] to advise us of your accommodation request. In your email, please include a description of the specific accommodation you are requesting as well as the job title and requisition number of the position for which you are applying.

 

Frequently Asked Questions

Is the salary disclosed for the Storage Rack Infrastructure Automation & Cluster Bring-Up - Hive Program position at sandisk?
The salary for this Storage Rack Infrastructure Automation & Cluster Bring-Up - Hive Program role at sandisk is not publicly listed. Click "Apply Now" to learn more about the compensation package on their official careers page.
Where is the Storage Rack Infrastructure Automation & Cluster Bring-Up - Hive Program position at sandisk located?
This Storage Rack Infrastructure Automation & Cluster Bring-Up - Hive Program role at sandisk is based in Center District, il, Kfar Saba, Kfar Saba, Center District, Israel. The position is listed as on-site or hybrid. Check the full job description or apply directly to confirm the work arrangement.
Is the Storage Rack Infrastructure Automation & Cluster Bring-Up - Hive Program role at sandisk full-time or part-time?
This is listed as a Full time position. It is posted as a Storage Rack Infrastructure Automation & Cluster Bring-Up - Hive Program role at sandisk.
How do I apply for the Storage Rack Infrastructure Automation & Cluster Bring-Up - Hive Program position at sandisk?
Click the "Apply Now" button on this page. You will be redirected to sandisk's official application portal hosted on smartrecruiters where you can submit your application directly.
When was the Storage Rack Infrastructure Automation & Cluster Bring-Up - Hive Program job at sandisk posted?
This Storage Rack Infrastructure Automation & Cluster Bring-Up - Hive Program position at sandisk was posted on Aug 27, 2026. Apply as soon as possible — early applications are often reviewed first.
Storage Rack Infrastructure Automation & Cluster Bring-Up - Hive Program
sandisk
Apply for this role ↗

You'll be redirected to sandisk's official application page on SmartRecruiters.