Apple is considering offering an AI server built around "M8" series chips and has discussed incorporating Nvidia networking hardware, The Information reports. The system would be sold to outside customers, potentially bringing Apple back into a business it left behind when it discontinued Xserve in 2011. Apple is reportedly targeting companies that want to run AI models on their own equipment, with a particular focus on inference and generating responses from trained models.
Trust me, consumers aren’t the ones buying 5k€ displays, nor the nearly 20k€ maxed out Mac Studios.
Feels like opposite of “enterprise friendly priced hardware”
We don’t know what the pricing will be.
Something to consider: All those nVidia GPU servers have discrete GPUs, meaning they need to have both RAM and VRAM, and any time there’s need to transfer something between RAM and VRAM, that’s overhead. Apple runs everything in a shared pool of memory - which at present is significantly slower than the VRAM on those nVidia GPUs since it’s DDR rather than HBM, but if they do what they did with the Ultra line of chips and go even further, e.g glue together 8 chips instead of 2, they might make up a bit of the difference in memory bandwidth. Or they could add HBM to these chips I guess.
And a single nVidia DGX B300 with 2.1 TB VRAM is several hundred thousand. And a single one of those is not enough to run the biggest models. By offering less powerful chips and cheaper memory, they might theoretically be able to offer tons of memory and acceptable, though lower, performance for much less money. Allows security-conscious companies to run high-end LLMs locally. And since they deal in shared memory, you might be able to use CPU instead of GPU if that’s more efficient in some specific part of inference, without transferring things between different memory spaces.
The macOS license agreement specifies you may only run it on Apple hardware. That means Xserve. So if you’re developing for macOS or iOS and want to use Xcode and target Mac platforms, you need this.
Or if you want to run server-size virtualization. You can install ESXi or another hypervisor on an Xserve and be compliant.
Apple hardware is one of the best platforms for low scale AI workloads. A simple M4 Pro MacBook Pro with 64GB RAM can comfortably run the best “at-home” model currently available (Qwen3-Coder-Next) around 30-40tkps.
A comparably specced AMD Ryzen AI 395+ can do about half of that.
If Apple delivered that in server-scale, easily scalable packages, there would definitely be buyers.
I don’t know man, anything that becomes a full-sized server/hardware has different pricing than mini-PCs. We are talking about a collaboration with Nvidia, a company that manages to overcharge everyone.
Of course the pricing will be different. My point here was that if Apple can jam that much compute into something as small as a Mac Mini, imagine what they can deliver at 4-5U scale.
Apple and Nvidia hardware combined might just be the solution to more scalable AI compute that uses less power.
i am assuming the performance comes from the soldered components improving signal integrity? the same reason the framework desktop has soldered memory? if true, then there would not be much customization, i imagine.
Why would anyone ever buy a server from apple? Isn’t it a consumer brand? Feels like opposite of “enterprise friendly priced hardware”
Well OpenAI allegedly bought a boatload of mac minis. And I dont think it was for brushing up their resumes and watching porn.
So, building Codex for minis?
Trust me, consumers aren’t the ones buying 5k€ displays, nor the nearly 20k€ maxed out Mac Studios.
We don’t know what the pricing will be.
Something to consider: All those nVidia GPU servers have discrete GPUs, meaning they need to have both RAM and VRAM, and any time there’s need to transfer something between RAM and VRAM, that’s overhead. Apple runs everything in a shared pool of memory - which at present is significantly slower than the VRAM on those nVidia GPUs since it’s DDR rather than HBM, but if they do what they did with the Ultra line of chips and go even further, e.g glue together 8 chips instead of 2, they might make up a bit of the difference in memory bandwidth. Or they could add HBM to these chips I guess.
And a single nVidia DGX B300 with 2.1 TB VRAM is several hundred thousand. And a single one of those is not enough to run the biggest models. By offering less powerful chips and cheaper memory, they might theoretically be able to offer tons of memory and acceptable, though lower, performance for much less money. Allows security-conscious companies to run high-end LLMs locally. And since they deal in shared memory, you might be able to use CPU instead of GPU if that’s more efficient in some specific part of inference, without transferring things between different memory spaces.
The macOS license agreement specifies you may only run it on Apple hardware. That means Xserve. So if you’re developing for macOS or iOS and want to use Xcode and target Mac platforms, you need this.
Or if you want to run server-size virtualization. You can install ESXi or another hypervisor on an Xserve and be compliant.
Apple hardware is one of the best platforms for low scale AI workloads. A simple M4 Pro MacBook Pro with 64GB RAM can comfortably run the best “at-home” model currently available (Qwen3-Coder-Next) around 30-40tkps.
A comparably specced AMD Ryzen AI 395+ can do about half of that.
If Apple delivered that in server-scale, easily scalable packages, there would definitely be buyers.
I don’t know man, anything that becomes a full-sized server/hardware has different pricing than mini-PCs. We are talking about a collaboration with Nvidia, a company that manages to overcharge everyone.
Of course the pricing will be different. My point here was that if Apple can jam that much compute into something as small as a Mac Mini, imagine what they can deliver at 4-5U scale.
Apple and Nvidia hardware combined might just be the solution to more scalable AI compute that uses less power.
I’m imagining someone sticking Mac Minis into a UCS chassis as the server blades and giggling a bit
i am assuming the performance comes from the soldered components improving signal integrity? the same reason the framework desktop has soldered memory? if true, then there would not be much customization, i imagine.