Bit errors occur at the physical layer, and packet loss is reflected in interface forwarding statistics. The two often appear at the same time but have completely different attribution paths. The first thing the operation and maintenance system needs to do is to have all teams use the same set of counters to report faults, and then peel them down hierarchically.

First align the counter caliber to a table

Commonly used indicators include input error frames, CRC check errors, output discard count and optical port alarm seconds. To capture a phenomenon, it is necessary to leave the device readings at both ends, timestamps, and the type of services carried by the link at the same time, because retransmission by upper-layer applications will conceal the true location of packet loss. The server side should also check the network card driver statistics, for examplePCI-EX8 25G dual optical port fiber optic network card (Intel E810 chip)Whether the packet loss count increases synchronously with the switch port. If the two sides are not consistent, it means that there is an aggregation link in the middle that has not been checked. It is recommended that this comparison table be posted at the door of the computer room and everyone should have a copy.

Peel it down in three layers: surface, segment and point.

  • Surface: Packet loss occurs on multiple devices on the same side, pointing to the backbone link, aggregation port, or abnormal power supply in the computer room.
  • Section: A single link continues to have bit errors. Focus on checking the cleanliness of the end face, the bending radius of the jumper and the light receiving margin.
  • Point: Only a certain host or a certain group of addresses is abnormal, mostly because the driver, duplex or speed gear does not match.
  • Intermittent faults require an additional round of long-term packet capture. Short-term observation can easily misjudge the problem as normal.

The self-loop is in the front and the replacement part is in the back

A common approach encountered by EB-LINK technical support on site is to replace the module first, but the more time-saving sequence is to self-loop first: use a short jumper to self-loop the port on the same side. If the error count stops increasing, it means that the equipment side is normal, and the fault is in the optical cable or distribution frame. If there are still errors in the self-loop, then check the matching of the port, module and written code. The replaced modules are labeled with serial numbers and symptoms, and are returned to the warehouse for unified retest before being judged as good or bad, to avoid scrapping usable parts as faulty parts. The entire round of demarcation process is written into the trouble ticket and signed by both parties at the handover, so that the next team does not have to do it all over again.