How can we help?

Understanding and Troubleshooting Packet Error Alerts in Auvik

Follow

Packet error alerts indicate that an interface is reporting an elevated number of malformed or damaged frames. Common examples include CRC, FCS, alignment, and too-long or too-short errors.

These alerts usually point to a problem on the physical link, an interface configuration mismatch, faulty hardware, or an overloaded device. They should not be dismissed automatically: even intermittent errors can cause retransmissions, poor application performance, and unstable links.

What the alert means

Auvik’s packet-error alert is intended to notify when packet errors exceed the configured threshold. The legacy preconfigured alert evaluates the number of packet errors over a five-minute period and clears when the count returns to or below the configured threshold for five minutes. Its default repetitive-alert behavior pauses the alert after five occurrences within two hours against the same entity.

In Alerts v2, the preconfigured High Packet Error(s) alert evaluates Ethernet interfaces and, by default, combines packet errors with interface utilization. The current documented default is more than 100,000 total packet errors, interface utilization greater than 10%, and a sustained evaluation period before notification. The clear condition uses a lower packet-error value to provide a buffer.

The exact behavior depends on whether the alert is a legacy definition or an Alerts v2 definition and whether an administrator has changed the defaults. Check the alert source, trigger condition, clear condition, and delay before interpreting the frequency of alerts.

Packet errors versus packet discards

Packet errors and packet discards are different conditions:

  • Packet errors usually indicate that a frame was received or transmitted with a physical or framing problem, such as a CRC, FCS, or alignment error.
  • Packet discards generally indicate that the device intentionally dropped traffic because of congestion, queue exhaustion, ACL or security policy, spanning-tree behavior, storm control, or another forwarding decision.

Use the packet-error troubleshooting path for CRC, FCS, alignment, and framing problems. Use the packet-discard path when the evidence points to congestion, queueing, policy, or forwarding behavior.

Why new or frequent alerts may appear

A recent physical or traffic change exposed an existing problem

Traffic may have increased enough to expose a marginal cable, optic, transceiver, interface, or port. A link that appeared healthy at low utilization can begin accumulating errors when it carries more traffic.

Review recent changes such as:

  • New devices, VLANs, trunks, or uplinks
  • Topology changes or link reconfiguration
  • Firmware or operating-system upgrades
  • Optic, cable, patch-panel, or transceiver replacement
  • Changes to speed, duplex, MTU, or negotiation settings
  • Increased traffic or a new high-throughput application

The alert definition became more sensitive

An administrator may have lowered the packet-error threshold, removed a utilization qualifier, reduced the alert delay, or added more interfaces to the alert scope. In Alerts v2, trigger conditions can include interface packet-error count, utilization, AND/OR logic, and time-based conditions, so a small definition change can materially increase alert volume.

A link moved from low to high utilization

Higher utilization increases the number of frames exposed to a marginal physical link and can make the same error percentage produce a larger absolute error count. Compare packet errors with interface utilization and traffic volume rather than looking at the error count in isolation.

The alert is reporting different entities than expected

Legacy alerts can generate separate notifications for individual interfaces. Alerts v2 can evaluate any or all interfaces on a device and consolidate the result into a device-level notification. A migration or scope change can therefore make alert volume appear to change even when the underlying network behavior is similar.

First-response checklist

When a packet-error alert triggers:

  1. Identify the device, interface, direction, error type if available, and alert source.
  2. Determine whether the alert is legacy or Alerts v2.
  3. Record the trigger threshold, evaluation period, alert delay, clear threshold, and current interface utilization.
  4. Check whether the error counters are increasing or whether the alert was caused by a historical counter crossing a threshold.
  5. Check the same interface at the remote end of the link.
  6. Correlate the alert timestamp with interface flaps, device logs, topology changes, and user or application impact.

What to check

1. Inspect interface counters

Use the device CLI, SNMP data, or Auvik’s interface statistics to determine whether counters are actively increasing. Look for:

  • CRC errors
  • FCS errors
  • Alignment errors
  • Runts, giants, or too-long/too-short frames
  • Input and output errors
  • Carrier or symbol errors
  • Interface resets or flaps

A single non-zero historical counter is less conclusive than a counter that increases steadily during the alert window. Capture the counter values, timestamps, interface name, and link partner before clearing or resetting counters.

2. Check the physical layer

Inspect both ends of the connection and the path between them:

  • Replace suspect patch cables.
  • Reseat or replace fiber or copper transceivers.
  • Inspect patch panels and intermediate connections.
  • Check for damaged, bent, contaminated, or unsupported fiber.
  • Confirm that the cable category and optic type match the interface.
  • Check for excessive distance, attenuation, or environmental conditions.

If the errors move with the cable or optic, the component is suspect. If the errors stay with the port, investigate the interface or hardware.

3. Verify speed and duplex

Confirm that both ends of the link negotiate the same speed and duplex mode. If the design requires hard-coded settings, configure both ends consistently. A mismatch—especially a full-duplex interface connected to a half-duplex interface—can cause collisions, late collisions, alignment errors, and poor performance.

Also verify that the interface is not unexpectedly falling back to a lower speed or changing negotiation state.

4. Review device and interface health

Check device logs, interface event history, and hardware diagnostics for:

  • Link up/down flaps
  • Transceiver alarms
  • Port or ASIC errors
  • Power or temperature warnings
  • Hardware module failures
  • Recent reloads or firmware events

If multiple interfaces on the same device begin reporting errors at the same time, consider a shared hardware, power, backplane, software, or environmental issue.

5. Correlate with traffic and configuration changes

Compare the alert with utilization, traffic direction, and recent changes. Determine whether the errors are isolated to one traffic direction or appear on both ends of the link. Review new broadcasts, multicast traffic, trunk changes, MTU changes, or large transfers that may have exposed a weak link.

Using Auvik to investigate

Review the interface alert definition

Open Manage Alerts and inspect the packet-error alert definition. Confirm:

  • Whether it is a legacy alert or an Alerts v2 alert
  • Which devices, interfaces, or tags are in scope
  • Whether it evaluates input, output, or total packet errors
  • The packet-error threshold
  • Any interface-utilization condition
  • The trigger delay or evaluation period
  • The clear threshold and clear delay
  • The notification channels

Auvik supports interface packet-error count as a trigger condition, and Alerts v2 can combine interface conditions with AND/OR logic and timing requirements.

Use alert variables in notifications

For interface alerts, include variables such as the interface packet-error count, interface name, interface description, and device name in the trigger message. This gives the responder enough context to start troubleshooting without first opening the Auvik dashboard.

Check whether the alert is already resolved

A transient error burst may clear when the error rate remains below the clear threshold for the configured clear window. A resolved alert is still useful evidence: compare its timestamp with device counters and logs to determine whether the problem was a one-time event or an intermittent fault.

Reducing noise without hiding real problems

Tune thresholds carefully

If alerts are too sensitive, adjust the threshold or evaluation period based on observed baseline behavior. Do not simply disable the alert or raise the threshold until genuine link faults disappear.

Use a baseline from several normal operating periods, including busy periods. A useful threshold should distinguish normal counter growth from abnormal error growth.

Separate critical links from access ports

Use stricter alerting for core, distribution, trunk, WAN, storage, and other critical links. Access ports may need a different threshold or delay because endpoints can be intermittently connected or physically disturbed.

If different link classes require different settings, create separate alert definitions or scopes rather than forcing one threshold across every interface.

Use an alert delay or sustained condition

Keep a sustained evaluation period or alert delay when short-lived link events are common. This allows transient blips to clear without generating a persistent notification while still catching errors that continue long enough to affect service.

Be aware that the legacy packet-error alert’s five-minute trigger and clear periods are fixed in the preconfigured definition. If you need different timing behavior, determine whether an Alerts v2 equivalent can provide the required time-based condition.

Keep a clear threshold buffer

When supported, set the clear threshold below the trigger threshold. This prevents the alert from repeatedly opening and closing while the error rate fluctuates near the boundary.

Avoid duplicate alert definitions

If both a legacy packet-error alert and an Alerts v2 packet-error alert are enabled with the same notification channel, responders may receive duplicate notifications. Compare the alert source and notification settings before enabling a second definition.

When to escalate

Escalate to the network or hardware team when:

  • Error counters continue rising rapidly after cables, optics, and patching are checked.
  • Errors occur on both ends of the link or move with the link component.
  • Multiple interfaces on the same device show errors simultaneously.
  • Interface flaps, hardware warnings, temperature events, or power issues are present.
  • Errors correlate with user-visible performance degradation, packet loss, voice or video quality problems, or application timeouts.
  • The issue persists after a controlled interface, optic, cable, or port substitution.

Escalate to the vendor when the evidence points to a failed port, transceiver, line card, backplane, or known firmware defect. Include timestamps, device and interface identifiers, counter samples, logs, optic details, and the Auvik alert definition.

Suggested incident record

Capture the following information for repeatability:

FieldExample
Device and interfaceSW-CORE-01 / Gi1/0/24
Link partnerSW-DIST-02 / Gi1/0/48
Alert sourceAlerts v2
Error typeCRC and FCS
Trigger and clear thresholds100,000 / 90,000
Evaluation or delaySustained for 3 hours
Utilization at trigger68%
Counter samples10:00 = 2,100; 10:15 = 48,900
Physical checksCable replaced; optic reseated
Device logsThree link flaps at 10:07
ResolutionMoved link to replacement optic

Summary

Packet-error alerts are signals of framing or physical-link problems, not merely a measure of high traffic. Start by confirming the alert definition and source, then determine whether counters are actively increasing. Check the physical path, speed and duplex, device health, and recent traffic or configuration changes before tuning the alert.

Use thresholds, sustained evaluation periods, clear buffers, and separate scopes for critical links to reduce noise safely. Escalate when errors persist, affect multiple interfaces, or correlate with service degradation.

Related Auvik knowledge-base articles

Was this article helpful?
0 out of 0 found this helpful
Have more questions? Submit a request

Auvik System Status

Check system status