Majestic Labs, a startup founded in 2023 by engineers who previously worked at Google and Meta, rolled out a new AI server it calls Prometheus. Instead of relying on traditional graphics processing units, the system uses up to 12 Ignite AI Processing Units (IPUs) per chassis. Each IPU blends Arm cores with RISC‑V vector and tensor engines, and the units share a single pool of LPDDR6 memory ranging from 8 TB to 128 TB. That memory, common in smartphones, replaces the costly high‑bandwidth memory (HBM) found on most GPUs.

The company’s pitch centers on a single technical obstacle: the memory wall. In many inference workloads, developers hit a ceiling not because the compute chips are too slow, but because they cannot reach enough fast memory. Majestic’s engineers answered that by decoupling memory from the processor and aggregating it through custom chiplets linked by copper cables up to a metre long. The result, they say, is a coherent memory pool far larger than any GPU‑centric design can address.

Majestic compares Prometheus to Nvidia’s DGX B300, which houses eight Blackwell GPUs and 2.3 TB of HBM. According to the startup, a single Prometheus rack delivers more than 50 times that amount of fast memory while providing 1.7 times the interconnect bandwidth. The same rack, Majestic asserts, can match the fast‑memory capacity of 25 Nvidia Vera Rubin racks, all while consuming a fraction of the power.

The hardware story is matched by a software approach that aims to keep the transition painless for developers. Prometheus supports open standards such as PyTorch, vLLM and OpenAI’s Triton, and Majestic claims that models built for GPUs can run on its platform without modification. The company says it has already secured orders from large enterprises, cloud‑native providers and hyperscalers, though no unit has left the factory.

Majestic’s claims remain unverified by independent testing. All performance figures come from the company itself, and the server has not yet shipped. The design also raises practical questions. A 128 TB memory pool built from 2 GB LPDDR6 dies would require roughly 64,000 individual chips, implying more than a hundred aggregation chiplets in a single server. Maintaining coherence across that many components is a non‑trivial engineering challenge.

Despite the uncertainties, the startup has attracted significant financial backing. Late last year, Majestic closed a $100 million funding round, a modest sum compared with the multi‑billion‑dollar budgets of its larger rivals. The firm now employs about 40 people split between Tel Aviv and Los Angeles, and it plans to begin shipping hardware in the coming year.

Majestic joins a growing roster of AI‑hardware startups that are trying to chip away at Nvidia’s dominance from unconventional angles. Competitors are exploring optical interconnects, edge‑focused inference silicon, and open‑networking gear. Each approach targets a different perceived weakness in Nvidia’s stack, whether it be power consumption, latency, or, in Majestic’s case, memory scarcity.

When Prometheus finally reaches customers, independent benchmarks will determine whether the memory‑first architecture lives up to its promise. Until then, the startup’s bold claims stand as a preview of a possible shift in how AI workloads are built and scaled.

Este artigo foi escrito com a assistência de IA.
News Factory APP - notícias agênticas para impulsionar seu SEO e AEO.