Troubleshooting Upgrade Failures
Upgrading a SmartZone cluster may
lead to various issues if the environment is not properly prepared. The following
checklist
outlines critical pre-upgrade steps and verification tasks designed to minimize the
risk of
upgrade failures.
- Supported AP Models: Review the release notes to ensure that all your Access Points (APs) are supported in the new release.
- Upgrade Path: Verify the upgrade path on the RUCKUS Support site to ensure the current version compatibility towards the destination version.
- Backup Configuration: Take cluster and configuration backup files from the existing cluster and save them to your local PC.
- Resource Requirements: For Virtual SmartZone setups, resource requirements may vary in newer releases due to additional features. Confirm the necessary resources in the RUCKUS SmartZone Upgrade Guide.
- MD5 Checksum Verification: Verify the MD5 checksum of the downloaded firmware file to ensure it is not corrupted.
- Active Support Contract: Ensure you have an active support license to perform the upgrade.
- Unsupported AP Firmware: Verify if the Zones have any unsupported AP firmware versions for the destination version of the controller.
- Cluster Upgrade Process: The multiple-node cluster upgrade process is similar to a single-node cluster upgrade. Initiate the upgrade from the Leader node, which will first upgrade the Follower node and then itself.
- AP Zone Upgrade: After upgrading the cluster, manually upgrade the AP zones to the required AP version.
- AP Functionality During Upgrade: During the cluster upgrade process, APs should function normally except for Tunnel WLANs, where clients will not be able to pass traffic through the controllers.
If the upgrade process fails, complete the following steps and then contact RUCKUS Support:
- For Virtual SmartZone setups, ensure
the resources for the virtual machines are correct on each node using the controller
CLI commands
show cpuinfo,show diskinfo, andshow meminfo. (Note: On hardware platforms SZ-144 and SZ-300, there is no need to check resources.) - Verify the services are up and
running on each node by executing the controller CLI commands
show serviceandshow cluster-state. - Ensure the NTP server is reachable and synchronized on all nodes. Check NTP details on the controller GUI under Administration > System > Time.
- Confirm there is no latency on the communication path to upload a new image file and between the cluster nodes. Enter the ping command in the CLI of the controller nodes.
- Take screenshots of all errors seen in the controller web UI during the upgrade.
- Note the timestamp of the failure.
- Enable the Debug level on Configurer and Web applications and replicate the upgrade.
- Collect Snapshot logs from each node for further investigation with RUCKUS Support.