Farm: partial outage

Incident Report for Farm HPC cluster

Resolved

All nodes and systems have been brought back into service.
Posted Aug 27, 2026 - 14:27 PDT

Update

The reboot of Farm's login node has allowed homedirs that were previously jammed to work again. Admins are still reviewing power usage before bringing nodes back online.
Posted Aug 27, 2026 - 09:57 PDT

Update

Farm's login node has not recovered its NFS mounts and will be rebooted at 9:40am. Please save your work and log out.
Posted Aug 27, 2026 - 09:33 PDT

Update

Most NAS have been brought back into service. Nodes will remain down until we can reaccess power tomorrow.
Posted Aug 26, 2026 - 19:13 PDT

Update

We are continuing to work on a fix for this issue.
Posted Aug 26, 2026 - 18:30 PDT

Identified

One of the power breakers that powers parts of both Hive and Farm tripped. Data Center operators are calling facilities.
Posted Aug 26, 2026 - 17:18 PDT

Update

We are continuing to investigate this issue.
Posted Aug 26, 2026 - 17:05 PDT

Investigating

Monitoring has alerted us that a part of Farm is currently offline. This includes nas-4-0, nas-4-1, nas-4-2, nas-4-3, nas-5-2 and nas-5-3. Admins are onsite and investigating.
Posted Aug 26, 2026 - 17:03 PDT
This incident affected: Login, Storage, high,low, bmh,bmm, gpuh,gpum and Virtualization (Proxmox Virtualization Nodes).