Hard Drive Failures
RustFS ensures read/write access can still be provided when some disks fail through mechanisms similar to erasure coding, and automatically heals data after replacing disks.
RustFS ensures read/write access can still be provided when some disks fail through mechanisms similar to erasure coding, and automatically heals data after replacing disks.
Table of Contents
- Unmount Failed Disk
- Replace Failed Disk
- Update
/etc/fstabor RustFS Configuration - Remount New Disk
- Trigger and Monitor Data Healing
- Follow-up Checks and Notes
Unmount Failed Disk
Before replacing the physical hard drive, you need to safely unmount the failed disk from the operating system level to avoid I/O errors in the file system or RustFS during the replacement process.
# Assume the failed disk is /dev/sdb
umount /dev/sdbNotes
- If there are multiple mount points, execute
umountseparately. - If encountering "device is busy", you can first stop the RustFS service:
systemctl stop rustfsReplace Failed Disk
After physically replacing the failed disk, you need to partition and format the new disk, and apply the same label as the original disk.
# Format as XFS and apply the same label as the original disk (RustFS requires XFS, like the other disks in the deployment)
mkfs.xfs -L DISK1 /dev/sdbRequirements
- New disk capacity ≥ original disk capacity;
- File system type consistent with other disks (XFS, per the installation guide);
- Recommend using labels (LABEL) or UUID for mounting to ensure disk order is not affected by system restarts.
Update /etc/fstab or RustFS Configuration
Confirm that the mount item labels or UUIDs in /etc/fstab point to the new disk. The mount point must stay identical to the one listed in RUSTFS_VOLUMES (in /etc/default/rustfs), so the healed disk rejoins the cluster at the same path.
# View current fstab
cat /etc/fstab
# Example fstab entry (no modification needed if labels are the same)
LABEL=DISK1 /data/rustfs0 xfs defaults,noatime 0 2Tips
- If using UUID:
blkid /dev/sdb
# Get the new partition's UUID, then replace the corresponding field in fstab- After modifying fstab, be sure to validate syntax:
mount -a # If no errors, configuration is correctRemount New Disk
Execute the following commands to batch mount all disks and start the RustFS service:
mount -a
systemctl start rustfsConfirm all disks are mounted normally:
df -h | grep /data/rustfsIf some mounts fail, please check if fstab entries are consistent with disk labels/UUIDs.
Trigger and Monitor Data Healing
Once RustFS detects a freshly formatted disk at a known mount point, its background scanner heals the missing data onto it automatically — no manual command is required. Confirm that recovery has started and track its progress through the service logs:
# For systemd-managed installations
journalctl -u rustfs -f
# Or view the log files under the directory set by RUSTFS_OBS_LOG_DIRECTORY
tail -f /var/logs/rustfs/rustfs.logYou can also open the RustFS Console and check the disk status of the affected node.
Notes
- The healing process will complete in the background, usually with minimal impact on online access;
- After healing is complete, the tool will report success or list failed objects.
Follow-up Checks and Notes
- Performance Monitoring
- I/O may fluctuate slightly during healing, recommend monitoring disk and network load.
- Batch Failures
- If multiple failures occur in the same batch of disks, consider more frequent hardware inspections.
- Regular Drills
- Regularly simulate disk failure drills to ensure team familiarity with recovery processes.
- Maintenance Windows
- When failure rates are high, arrange dedicated maintenance windows to speed up replacement and healing.