Back to feed

Meta and Panmnesia propose CXL architecture that could support up to 960 AI accelerators

2 min
Meta and Panmnesia propose CXL architecture that could support up to 960 AI accelerators

This digest was compiled by AI from multiple sources — links to the originals are below.

Meta and Panmnesia have proposed a Compute Express Link (CXL) architecture to connect up to 960 AI accelerators in a single coherent data center domain. The design aims to reduce communication delays between racks, replacing conventional Ethernet or InfiniBand networks. The architecture increases accelerator coordination from two to sixteen devices per CPU.

Key Facts

  • The proposed CXL architecture could bring as many as 960 AI accelerators into one coherence domain.
  • One CPU could coordinate 16 accelerators under Panmnesia's design, an eightfold increase over NVIDIA's GB200 NVL72 reference configuration.
  • Panmnesia's fabric controller and link acceleration unit have completed silicon validation, and its switch has been fabricated.
  • The architecture uses a high-fan-out switch, a link acceleration unit, and a fabric controller organized into trays, pods, and a fabric.

CXL Architecture

The design uses Compute Express Link (CXL) to connect CPUs, accelerators, and memory across multiple racks without conventional network links. CXL provides a shared coherence mechanism, allowing processors, accelerators, and memory to participate within one connected resource environment. The proposed architecture adds dedicated hardware to keep communication paths and processing behaviour more consistent across the larger fabric. Panmnesia's design uses a high-fan-out switch, a link acceleration unit, and a fabric controller to manage traffic across the system. These components are organized into trays, pods, and a fabric, borrowing organizational principles normally associated with arranging functional blocks inside semiconductor chips.

Accelerator Coordination

The review compares the proposed arrangement with NVIDIA's GB200 NVL72, where one CPU directly coordinates two accelerators through NVLink-C2C. Under Panmnesia's architecture, one CPU could coordinate 16 accelerators, representing an eightfold increase over that reference configuration. According to the published design, around 60 such groups could then form a coherence domain containing approximately 960 accelerators.

Silicon Validation

The company says its fabric controller and link acceleration unit have completed silicon validation. Its switch has also been fabricated, while pre-release silicon is reportedly being supplied as development continues toward commercial products.

1 source

Time · lag behind first