HMON Overview
Hosted applications and processes that are registered with Health Monitor in the FastIron operating system are monitored as high-availability (HA) processes. Health Monitor (HMON) provides the following services on ICX devices for registered applications and processes:
- Monitoring at specified intervals and restarting the process on failure
- On-demand process start or stop
- Process stop or start based on the device role in a stack
HMON Process Registration
For each process registered with HMON (also referred to as client processes), a set of attributes can be configured. These attributes include the following parameters:
- Monitoring interval - specified in 10 second increments
- Process criticality - critical or non-critical
- Maximum number of restarts - configurable cap on restart and recovery
- Stack role mask - the stack roles under which the process can run
Dynamic Start and Stop
When a process is registered with HMON as an "on demand" process, FastIron sends IPC messages to HMON to start and stop the process dynamically based on defined functionality. HMON monitors the process from the time it is started until it is stopped.
Process Availability Based on Stack Role
Some processes are dependent on the current role of the ICX device in a stack configuration. For example, a particular application may run on the device only while the device is acting as the active controller. Role-dependency can be configured as part of HMON client registration, and the client process can be started and stopped in relation to the role.
Clients Marked as Faulty
When a client process exceeds the maximum number of recovery attempts, it is marked as faulty and is no longer available. The failure is logged as a syslog message. Cyclic crashes are an indication of an issue that should be addressed. Contact Ruckus technical support for assistance.
Critical Processes
If the faulty client has been registered with HMON as system critical, the ICX device reboots. In a stack configuration, the standby controller takes over as the active controller, which may restore the client process to normal operation.
To determine whether a registered HMON client is a
critical or non-critical process, enter the show hmon client configuration
all-clients command in Privileged EXEC mode. Refer to Troubleshooting
HMON for an example of show hmon client configuration
all-clients command output.
Determining the Administrative and Operational State of an HMON Client
Enter the show hmon client status
all-clients command in Privileged EXEC mode to display both the
administrative and operational state of all HMON clients registered on the device.
Administrative states are defined as follows:
- Enabled, Not Started, HA Disabled - The client process is enabled; however, it is not started, as it is not qualified to run on the current stack role. As a result, the process is not being monitored.
- Enabled, Started, HA Enabled - The application is enabled, started or running, and monitored for HA.
- Disabled, HA Disabled - The application is not enabled, not started or running, and not monitored for HA.
The operational state provides additional information for each client process and can be one of the following:
- Up - The client process is up and running.
- Down - The client process is not running.
- Recovering - HMON has initiated a recovery for the crashed or failed process, and the process is recovering (transient state).
- Recovery Failed - Recovery has failed.
- Faulty - Due to repeated failed recoveries, the maximum allowable recovery attempts has been exceeded, and the client process is marked as faulty.
Dependent Process Restart
Beginning with FastIron release 09.0.00, when a process monitored by HMON restarts for recovery, the dependent processes also gets restarted automatically.
Dependent processes are defined as those processes that require a restart when another process is going for an HA restart recovery. All the dependent processes are grouped together and each process in the group have a criticality defined as following:- HMON_DEP_PRCSGRP_CRITICAL: Other processes in the group should be restarted if this process restarts.
- HMON_DEP_PRCSGRP_NON_CRITICAL: Other processes in the group should not be restarted if this process restarts.
A process group can have more than one HMON_DEP_PRCSGRP_CRITICAL member in it. Processes should be grouped for dependency only if there is at least one HMON_DEP_PRCSGRP_CRITICAL member in it.
The following syslog message is displayed when the dependent process is restarted:
Application <process name> is being restarted as the process <HMON_GROUP_CRITICAL process name> which it depends on got restarted
When process restart fails and HMON exceeds the maximum number of restart recovery attempts, the following critical syslog message is displayed to notify the unavailability of functionality provided by this process.
Application <process name> failed recovery and functionality provided by it will not be available until failure reason is remedied (may require manual intervention)Contact RUCKUS Technical Support if this type of failure occurs.
Refer to Troubleshooting
HMON for an example of show hmon client status
all-clients command output.