ISSU Errors

There are several sources of errors that may be encountered during an ISSU, and there are two means of error recovery.

Common Errors During ISSU

Error Description
Hot-swap timeout Unit hot-swap does not complete within the expected time.
Version synchronization timeout Version information synchronization does not complete within the expected time.
Standby assignment timeout After upgrading the current standby unit, the standby assignment does not occur within the expected time.
Standby assignment error After upgrading the current standby unit, the expected unit was not elected as the new standby unit.
Image/boot source mismatch After a unit upgrade, the image version and boot source did not match the expected version or boot source.
Unit fails to rejoin The unit fails to rejoin the stack within the specified time after an upgrade.
Unit delete The unit is detached from the stack while the ISSU is in progress.
Ping fail A unit fails to respond to keepalive messages.

Crash and Manual Abort Errors

Error Message Description
Unit crash If the issu primary command on-error option is specified, the unit that crashes is reloaded from the partition specified in the command.

The active controller detects this condition as a unit delete and reloads all the existing stack members from the partition specified in the issu primary command on-error option.

Active reload/crash If the active controller reloads unexpectedly, or crashes while the ISSU is in progress, the stack units detect the loss of the active controller and abort the ISSU.

If the issu primary command on-error option is specified, all units that were part of the stack at the time of the active controller crash are reloaded from the partition specified in the command. Any units that were being upgraded at the time of the active controller failover reload from the target partition given in the issu primary command.

Once all units have booted and an active controller has been elected, if some units have a running image different from the active controller image, an image auto-copy is executed, and units are reloaded to ensure they are all running the same image.

Manual abort If ISSU is aborted through the issu abort command, ISSU is stopped, and the stack is left in the current state for manual recovery.

This behavior occurs whether ISSU is started with or without the issu primary command on-error option.

Error Recovery

There are two means for error recovery, one manual and one automatic:

  • When ISSU is started with the issu primary or issu secondary command, the following results apply:

    • If an error occurs, the upgrade is aborted, and the stack is left for manual recovery. In this condition, it is likely that the running images on the stack units are different. After abort, image auto-copy is not executed.
    • Units continue with their current running image until the system is reloaded. As a result, a reload of the entire stack is required to bring it back to a functional state.
    • To ensure system stability, the stack is left in the aborted state. You must reload the system manually. If any of the stack units are reloaded individually, they cannot move to the Ready state. To execute a manual recovery, refer to Manual error recovery.
  • The following points apply when ISSU is started with an issu primary or an issu secondary command that includes an on-error reload-primary or an on-error reload-secondary option:

    • If an error occurs, the upgrade is aborted.
    • All the units in the stack are automatically reloaded to the partition specified by the issu primary command on-error option.
    • After the system reload, any units that were unreachable at the time of the ISSU abort may have an image that is different from the other units. When these units rejoin the stack, an image auto-copy is executed for any units with a mismatched image, and they are reloaded after the auto image copy completes.