CodingBox Documentation

AI networking overview

Training large AI models spreads work across many GPUs across many servers. The network between them — not any single GPU — often sets how fast a cluster can train, which is why AI fabrics lean on the highest-bandwidth optical transceivers available.

Why the network is critical

  • Collective operations — GPUs constantly exchange gradients and parameters in

all-to-all patterns, producing heavy east-west traffic.

  • Tail latency matters — a step waits for the slowest link, so consistent low

latency and lossless delivery are essential.

  • Scale — thousands of GPUs mean dense, high-radix switches and enormous optics counts.

Two fabric styles

  • InfiniBand — long established in HPC, with RDMA and lossless flow control (see

InfiniBand).

  • Ethernet with RoCE — RDMA over Converged Ethernet, increasingly used for AI at

scale on lossless Ethernet.

Optics

AI links are typically 400G and 800G, on QSFP-DD and OSFP modules. The specific interconnect choices and optics are covered in AI interconnects & optics.

The same form factors and management memory apply — CodingBox reads these high-rate modules like any other.