Cluster Nodes and High Availability

RUCKUS Network Director offers a comprehensive High Availability (HA) solution (DB replication) with near real-time database sync-up between RND nodes. HA clusters are used for failover, load balancing, and backup purposes.

RUCKUS Network Architecture

RUCKUS Network Director follows an Active-Active clustering model and the database synchronization is performed by the database itself, which takes the “Primary” and “Standby” roles. The data is replicated from “Primary” to “Standby” immediately when data is committed, which ensures the data stays in sync. This solution requires a pair of RND nodes to form a cluster. The pgpool, which is the middleware that communicates with the RND and database in both nodes, becomes the "Leader" and the "Follower". In case of failure, the nodes can fail over without any downtime. Because the current deployment requires the RND node to be in the same subnet, RND cannot be deployed across data centers. The RND deployment in each data center is independent and cannot communicate or share data with different regions.

HA Solution Data Flow

When the second node is added to form a cluster, pgpool in both nodes becomes the leader and the follower. If one node goes down, pgpool takes the responsibility to change the role of the database.

High Availability Workflow

The following example explains the data flow. If you access the system through Node 1 and create an AP registration rule, the data will be written to the database in Node 1. Simultaneously, the primary database in Node 1 replicates the data to the secondary database (as displayed in path 1 > 2 > 3 > 4 in the workflow illustrated in the figure).

Because the primary database is responsible for replicating the data, when you access the system through Node 2, the data is redirected to the primary database by the pgpool in Node 2, the data is saved in the Node 1 primary database, and replicated to the secondary database (as displayed in path 5 > 6 > 7 > 4 in the workflow illustrated in the figure). This approach ensures data consistency in each node.

If one of the RND nodes goes down, the other RND node takes over the responsibility to serve the system.

There are different scenarios:

  • If both nodes reboot consecutively in a very short time interval: The role of each node will not be changed. The pgpool retries several times (configurable) to detect if any other peer is available in the network. If it does not detect any peer, it will start to change its role and manage the database role accordingly.
  • If one of the RND nodes reboots: Basically no change, after node bootup, the pgsql will help to control themselves.
  • If one of the RND nodes is shut down: The other live node will change the role. For example, the secondary/follower will become the primary/leader for the pqsql and postgresql. Later, if the node rejoins, it will become the secondary/follower.