Troubleshooting SmartZone Node not Joining the Cluster

When integrating a new node into a SmartZone cluster, various issues can arise related to network reachability, firmware version consistency, form-factor compatibility, interface configuration, IP version alignment, VM resource levels, and latency requirements.

Review the following to learn more about troubleshooting these aspects and ensuring a smooth integration process.

  1. Verify Network Reachability: Ensure there is network reachability between the Leader node and the new node attempting to join the cluster. This can be confirmed by performing a ping test from the CLI of the controller nodes.
  2. Firmware Version Consistency: Confirm that the new node is operating on the same branch and build firmware version as the existing Cluster nodes.
  3. Form-Factor Compatibility: Ensure that the new node matches the form-factor of the cluster nodes. For instance, SmartZone hardware can form a cluster only with other SmartZone controllers of the same physical form-factor (for example, SZ144 with other SZ144 nodes, and SZ300 with other SZ300 nodes).
  4. Version Alignment for Virtual SmartZone: For Virtual SmartZone nodes, verify that both the connected node and the new node are using the same version, whether it is the High-Scale or Essentials version.
  5. Interface Configuration Matching: Ensure that the interface configuration of the new node aligns with that of the cluster nodes. For example, for Virtual SmartZone High Scale (vSZ-H), the configuration should match either the single-interface or three-interface mode of the connected nodes. For hardware-based SmartZone controllers, the configuration should match the single-port group or two-port group setup. A mismatch will result in the error: Error on checking cluster port group setting. Please make sure the port group setting are compatible. on the GUI. Enter the show interface command on the controller CLI to verify this.
  6. IP Version Consistency: The IP version (IPv4, IPv6, or Dual) should be consistent across all nodes.
  7. VM Resource Level Validation: Ensure that the VM resource levels are consistent between the nodes. A mismatch will trigger the error: “Error on checking node resource plan. The cluster requires all resource plan to be the same”.
  8. Latency Requirements: Verify that the latency between the SmartZone nodes is within acceptable limits, as high latency can prevent the node from joining the cluster. The latency requirements can be found among the network requirements in the Release Notes for the specific firmware version.
  9. Cluster Member Status: Confirm that all current members of the cluster are in a connected status and that their services are online. This can be verified by logging into the CLI of the cluster nodes and executing the following commands:
    • Execute the show service command to verify that all services are in an online status. Delete or attempt to reconnect any node that is not online in order to join the new node.
    • Execute the show cluster-state command to ensure that none of the following keywords appear in the system-state or cluster-state: Out of Service, Maintenance, Crash, Suspend, NetworkPartitionSuspected.
  10. Perform a factory reset: If you encounter the error “Error on first time initialization Process,” perform a factory reset on the new node. Ensure that only the new node is reset and not any active member of the cluster. After the reset, attempt to join the cluster again.

If all the previous troubleshooting steps have been attempted without success, please proceed with the following actions and open a support case with RUCKUS:

  • From the current Leader node, change the logging level of Web and Configurer applications to Debug level. Attempt to join the node again and then collect the Snapshot log. Once the logs are downloaded, revert the logging level back to Warning to avoid overwhelming the system resources. For detailed instructions on downloading Snapshot logs, refer to Viewing and Downloading Logs.
  • Additionally, collect the output of the show service and show cluster-state controller CLI commands from all connected nodes. This information will be crucial for the RUCKUS Support team to diagnose and resolve the issue effectively.