Troubleshooting HMON

Note: HMON commands provide output specific to the device on which the command is executed.
Note: Refer to the Ruckus FastIron Command Reference Guide for more information on the commands.

When an HMON client is marked as faulty, a syslog message similar to the following is issued:

Application webserver failed recovery and functionality provided by it may not be available until FastIron acts on it

Work with Ruckus technical support if this type of failure occurs. Before contacting technical support, perform the following steps to gather diagnostic information.

  1. In Privileged EXEC mode, enter the hmon status command and check the client list to confirm that the failed process is indeed being monitored by HMON. Capture the output.
    device# hmon status
    ----------------------
    Health Monitor Status: 
    ----------------------
    Hmon's Stack Role is : Standalone
    Number of Clients    : 4
    
    Client Names (ID) : 
    	nginx (4)
    	uwsgi-2.7 (5)
    	PySzAgtSrv.py (6)
    	dhcpd (3)
    
    The previous example of the hmon status command displays information for four registered HMON processes identified by their Client name and ID. The output also indicates that the device is not part of a stack.
  2. Enter the hmon client status all-clients command to check the administrative and operational state of the application. Capture the output.
    device# hmon client status all-clients 
    -----------------------------
    Health Monitor Client Status: 
    -----------------------------
    
    Status for client ID 4:
    Process Name              : nginx
    Valid                     : Yes
    Admin. State              : Enabled, Started, HA Enabled
    Oper. State               : Up
    
    Status for client ID 5:
    Process Name              : uwsgi-2.7
    Valid                     : Yes
    Admin. State              : Enabled, Started, HA Enabled
    Oper. State               : Up
    
    Status for client ID 6:
    Process Name              : PySzAgtSrv.py
    Valid                     : Yes
    Admin. State              : Enabled, Started, HA Enabled
    Oper. State               : Up
    
    Status for client ID 3:
    Process Name              : dhcpd
    Valid                     : Yes
    Admin. State              : Disabled, HA Disabled
    Oper. State               : Down
    
    The previous example of the hmon client status all-clients shows that the dhcpd process is disabled.
  3. Enter the show log command and check for hmond entries similar to those in the following example.
    Dynamic Log Buffer (4000 lines):
    Mar  4 22:22:19:E:hmond[392]: Client uwsgi-2.7 has reached/exceeded max funcmntr fail count: 2, initiating recovery from state UtilReportedFail  
    Mar  4 22:22:19:E:hmond[392]: Client uwsgi-2.7 is not functional, fail count is: 2, funcmntr fail count limit is: 2  
    Mar  4 22:22:09:E:hmond[392]: Client uwsgi-2.7 is not functional, fail count is: 1, funcmntr fail count limit is: 2  
    
  4. Enter the hmon client statistics all-clients command and capture the output.
    device# hmon client statistics all-clients 
    ---------------------------------
    Health Monitor Client Statistics: 
    ---------------------------------
    
    Statistics for client ID 4:
    Process Name                             : nginx
    Most recent PID                          : 1328
    Func. Monitor fail counts                : 0
    Total number of admin stops              : 0
    Total number of disallowed admin stops   : 0
    Total number of admin starts             : 1
    Total number of disallowed admin starts  : 1
    Total number of admin restarts           : 0
    Total number of restarts for recovery    : 0
    Code from latest func. fail recovery     : 0x0
    Status from latest Func. Monitor check   : Invalid
    
    Statistics for client ID 5:
    Process Name                             : uwsgi-2.7
    Most recent PID                          : 1343
    Func. Monitor fail counts                : 0
    Total number of admin stops              : 0
    Total number of disallowed admin stops   : 0
    Total number of admin starts             : 1
    Total number of disallowed admin starts  : 1
    Total number of admin restarts           : 0
    Total number of restarts for recovery    : 0
    Code from latest func. fail recovery     : 0x0
    Status from latest Func. Monitor check   : Invalid
    
    Statistics for client ID 6:
    Process Name                             : PySzAgtSrv.py
    Most recent PID                          : 1356
    Func. Monitor fail counts                : 0
    Total number of admin stops              : 0
    Total number of disallowed admin stops   : 0
    Total number of admin starts             : 1
    Total number of disallowed admin starts  : 1
    Total number of admin restarts           : 0
    Total number of restarts for recovery    : 0
    Code from latest func. fail recovery     : 0x0
    Status from latest Func. Monitor check   : Access Issue
    
    Statistics for client ID 3:
    Process Name                             : dhcpd
    Most recent PID                          : Not Available
    Func. Monitor fail counts                : 0
    Total number of admin stops              : 0
    Total number of disallowed admin stops   : 0
    Total number of admin starts             : 0
    Total number of disallowed admin starts  : 0
    Total number of admin restarts           : 0
    Total number of restarts for recovery    : 0
    Code from latest func. fail recovery     : 0x0
    Status from latest Func. Monitor check   : Not Invoked
    
    
  5. Enter the hmon client configuration all-clients command and capture the output.
    device# hmon client configuration all-clients 
    ------------------------------------
    Health Monitor Client Configuration: 
    ------------------------------------
    
    Configuration attributes for client ID 4:
    Process Name                  : nginx
    Startup Script                : nginx-service.sh
    Stackrole mask                : 0x3
    Starts on bootup              : No
    Process restartable           : Yes
    Criticality of the process    : Non-Critical
    Process restart count limit   : 5
    Heart-Beat monitoring reqd.   : No
    Functionality monitoring reqd.: Yes
    Func. monitoring interval     : 10 Secs
    Func. fail count limit        : 2
    
    Configuration attributes for client ID 5:
    Process Name                  : uwsgi-2.7
    Startup Script                : uwsgi-service.sh
    Stackrole mask                : 0x3
    Starts on bootup              : No
    Process restartable           : Yes
    Criticality of the process    : Non-Critical
    Process restart count limit   : 5
    Heart-Beat monitoring reqd.   : No
    Functionality monitoring reqd.: Yes
    Func. monitoring interval     : 10 Secs
    Func. fail count limit        : 2
    
    Configuration attributes for client ID 6:
    Process Name                  : PySzAgtSrv.py
    Startup Script                : pySzagent-service.sh
    Stackrole mask                : 0x3
    Starts on bootup              : No
    Process restartable           : Yes
    Criticality of the process    : Non-Critical
    Process restart count limit   : 5
    Heart-Beat monitoring reqd.   : No
    Functionality monitoring reqd.: Yes
    Func. monitoring interval     : 10 Secs
    Func. fail count limit        : 2
    
    Configuration attributes for client ID 3:
    Process Name                  : dhcpd
    Startup Script                : dhcpd-script.sh
    Stackrole mask                : 0x3
    Starts on bootup              : No
    Process restartable           : Yes
    Criticality of the process    : Non-Critical
    Process restart count limit   : 5
    Heart-Beat monitoring reqd.   : No
    Functionality monitoring reqd.: Yes
    Func. monitoring interval     : 10 Secs
    Func. fail count limit        : 2
    
    The previous example of the hmon client configuration all-clients command shows that all HMON clients running on the device are non-critical; that is, the processes become unavailable when marked faulty, and no unit reboot is attempted to initiate switchover to the standby controller.
  6. Enter the supportsave all followed by the IP address of the tftp server where supportsave logs are to be uploaded as shown in the following example.
    Note: For more supportsave command options, refer to the RUCKUS FastIron Command Reference.
    ICX7650-48P Router# supportsave all 10.22.141.59
    
  7. Collect the logs.
  8. Contact Ruckus technical support.