AI · 14 September 2026
CXL Architecture: Meta and Panmnesia Link 960 AI GPUs as One
Meta and South Korean startup Panmnesia have proposed a CXL-based data centre architecture that connects up to 960 AI accelerators across multiple racks as one coherent computing domain.
What happened
Meta has proposed a new data centre architecture, developed in partnership with South Korean startup Panmnesia, that uses Compute Express Link (CXL) technology to connect as many as 960 accelerators — effectively close to 1,000 AI GPUs — within a single coherent computing domain spanning multiple server racks.
The design aims to make a cluster of GPUs and other accelerators behave, from a memory and compute-coherence standpoint, like components on one large chip, rather than as separate machines linked by conventional networking. Panmnesia, a small specialist in CXL interconnect technology, is credited alongside Meta as co-author of the proposed architecture.
Why it matters
Large AI training and inference workloads are increasingly bottlenecked not by raw compute but by how efficiently GPUs can share memory and data across racks. A CXL-based architecture that pools hundreds of accelerators into one coherent domain points to a way of reducing the overhead and latency that come from treating each server as an isolated island, which could materially change how hyperscale AI infrastructure is designed and scaled.
For technology and infrastructure leaders, this signals that the next constraint on AI progress may be architectural — how compute, memory and interconnects are organised — rather than simply the number of chips deployed. Vendors and operators planning next-generation AI data centres will likely watch how such coherent-domain designs mature from proposal to deployment.
By the numbers
- 960 accelerators can reportedly be connected within a single coherent computing domain under the proposed architecture.
- Almost 1,000 AI GPUs per domain is the scale Meta and Panmnesia say the design can support.
- Multiple racks of servers are spanned by the single coherent domain, rather than being confined to one rack or server.
The Renascence take
It is tempting to read this purely as a hardware story, but it is really a lesson in how experience — whether for an end customer or an internal engineering team — is shaped by architecture decisions made years before a product ever reaches a user. The gap between "we have enough GPUs" and "our GPUs work coherently together" is the same gap that separates a company with lots of data from one that can actually act on it in real time.
Most organisations chase capacity — more compute, more headcount, more tools — when the real constraint is coherence: how well the parts already in place talk to each other. Meta and Panmnesia's approach is a reminder that scaling AI, like scaling service, is less about adding more units and more about removing the friction between them. Leaders investing in AI infrastructure should ask not "how many GPUs do we have" but "how efficiently can they act as one system" — because that question, more than raw capacity, will determine what AI can actually deliver to customers and employees downstream.
Sources
This briefing was written by our Newsdesk, synthesising reporting from the outlets below. Follow the links for the original coverage.
FAQ
Questions we get on this topic
More in AI
Stay ahead of CX
Get the signal, not the noise.
The stories shaping customer experience — plus the Journal and Experience Loom — in your inbox.