How can we help?

Troubleshooting intermittent port-down alerts on management controllers

Follow

What this issue looks like

You may receive repeated port-down or interface-down alerts even though the physical interface appears stable in the device or management-controller logs.

This can occur when Auvik temporarily loses access to a management controller, such as Cisco CIMC, HPE iLO, or Dell iDRAC. The controller may briefly stop responding to management requests even though the physical interface and the device itself remain operational.

Possible causes

Intermittent port-down alerts may be caused by:

  • Temporary loss of connectivity between the Auvik collector and the management controller.
  • Network latency, packet loss, WAN instability, VPN flaps, or firewall changes.
  • The management controller becoming temporarily unresponsive.
  • Management-controller CPU or control-plane load.
  • Management traffic being rate-limited or affected by control-plane protection.
  • Session or resource limits on the management controller.
  • A transient controller restart, sensor reset, or firmware-related event.
  • A monitored management IP address changing or becoming unreachable.

What to check

1. Confirm the alert and monitored entity

Identify:

  • The affected device and interface.
  • Whether the alert is a legacy alert or an Alerts v2 alert.
  • The monitored IP address or management interface.
  • The exact time and frequency of the alerts.
  • Whether multiple devices or interfaces alerted at the same time.

If multiple alerts occur together, look for a shared management-network, collector, firewall, WAN, or upstream-device issue.

2. Check the management-controller logs

Review the event logs for the affected controller and look for:

  • Controller restarts.
  • Temporary loss of network connectivity.
  • Sensor resets.
  • Management-service restarts.
  • Authentication or session errors.
  • CPU, memory, temperature, or hardware warnings.
  • Firmware or configuration events.

Compare the controller log timestamps with the Auvik alert timestamps.

3. Verify collector-to-controller connectivity

Confirm that the Auvik collector can consistently reach the monitored management IP.

Review:

  • Firewall and ACL rules.
  • Routing and VLAN changes.
  • VPN connectivity.
  • Management-network availability.
  • Any recent changes that could affect traffic from the collector.
  • Whether the controller responds consistently to the management protocols used by Auvik.

A device may remain operational while Auvik reports the interface as down if the collector temporarily cannot reach the monitored management endpoint.

4. Check for management-plane resource limits

If the controller logs show session, resource, or rate-limit events, have the device administrator review the controller configuration and vendor guidance.

Do not change controller session limits or control-plane protection settings without confirming the effect on the device’s security and management operations.

5. Review the alert timing

For Alerts v2, review the alert definition’s delay and trigger conditions.

A short alert delay may allow brief management-controller interruptions to generate notifications. A longer, carefully selected delay can prevent an alert from being created when the condition clears before the delay expires.

Use a delay that reduces transient noise without hiding a genuine outage.

For more information, see How to set Alert Delays with Alerts v2.

6. Review health-check settings

If the alert is based on device availability, review the applicable health-check frequency and minimum-failure settings.

Increasing the number of failures required before a device is considered offline may reduce alerts caused by short interruptions. Apply changes only to the affected device types or sites where appropriate.

For more information, see How do I manage device health check frequencies?.

Validate the result

After making a configuration change:

  1. Monitor the affected device during the period when alerts usually occur.
  2. Compare new alert timestamps with controller and device logs.
  3. Confirm whether the physical interface remains stable.
  4. Verify that the alert clears normally when connectivity returns.
  5. Confirm that a longer delay has not prevented notification of a genuine outage.

Avoid testing by intentionally disconnecting a production management controller unless the test is approved and performed during a maintenance window.

When to escalate

Escalate to the device or network administrator when:

  • Controller logs show restarts, sensor resets, resource exhaustion, or session-limit events.
  • Multiple devices report management-controller interruptions.
  • Firewall, VPN, routing, or ACL changes correlate with the alerts.
  • The controller remains reachable from other systems but not from the Auvik collector.
  • The physical interface is stable but the management controller continues to report intermittent failures.

Contact Auvik Support when:

  • The collector is online and network access has been verified, but Auvik continues to report intermittent failures.
  • The issue affects multiple unrelated devices or sites.
  • The alert behavior does not match the configured trigger, delay, or clear conditions.
  • Collector-level diagnostics or packet captures are required.

Include the following information when contacting Support:

  • Site and collector name.
  • Device and interface name.
  • Management-controller type.
  • Alert name and alerting platform.
  • Alert timestamps and frequency.
  • Monitored IP address.
  • Relevant controller and device log entries.
  • Recent network, firewall, VPN, or configuration changes.

Related articles

Was this article helpful?
0 out of 0 found this helpful
Have more questions? Submit a request

Auvik System Status

Check system status