As the video points out, there is a hardware/software feedback loop at play here. Since the hardware is deeply pipelined SIMD FPU datapath, the software is hand tuned machine coded transformations. In addition to limiting flexibility (every variation has to be ground out in CUDA), this prevents sparse matrix type optimisations. I predict that a set of general purpose CPUs - non shared memory at this scale - would be a much better use of transistor/power. And you can code that at high level give a CSP style approach such as https://bil-lang.org
I wonder what would "compute" mean if cpus were more efficient at matrix multiplication say 15 years ago. And on the flipside, what it would take to say train a frontier model entirely on cpus in the future.
On the compute side there's the issue of scalability. CPUs are designed to perform a handful of operations at one time. They typically have a small number of dedicated integer, float, and other ALU configurations. Having dedicated matrix multiplication instructions would still lock that to how many matrix-capable ALUs there are in the CPU.
GPUs are designed to process a large number of calculations at once (so they can process triangles in 3D graphics). This makes them good at ML applications as they can process many of the matrix calculations at once. A 4090 has 16,384 CUDA cores (general compute ALUs) and 512 Tensor cores (dedicated matrix compute ALUs); a 5090 has 21,760 CUDA and 680 Tensor cores.
The other issue when training models (and running larger models) is the amount of VRAM (or RAM for CPUs) available. GPUs are limited in this aspect, whereas CPUs can have a lot higher memory. This affects things like batch size and the size of model that can be trained or fine-tuned.
https://unsloth.ai has guides for how to fine-tune existing models like Qwen 3.8 27B, memory requirements, etc.
https://medium.com/@kailaspsudheer/the-transformers-arithmet... has some information on training a base model. A 7B llama model is estimated at taking ~34GB memory for inference at F32, but was observed requiring 96GB memory when training (for the model weights, gradients, activations, and optimizer states).
Note: you can reduce the memory required for training by recomputing the gradients, at a cost of performance/time. You can also do other tricks like performing a QLoRA/LoRA pass on the model then merging that into the model to create a checkpoint.
I don't know what sized model you could train on 64GB/128GB RAM via a CPU.
I wonder what "compute" would mean if CPUs were more efficient at matrix multiplication and vendors had the balls to pair each core to its own dedicated DDR and a star interconnect between.
Yes. there is real hardware architecture behind Apple’s “Unified Memory”, but the underlying idea is not uniquely Apple. What Apple did was design the CPU, GPU, memory controller, cache hierarchy, interconnect, package, OS, and graphics APIs together around the architecture. I mean, they can do shit like that because they own the entire product pipeline. They can fine tune the hardware in ways other OEMs can't.
In a Intel+Nvidia CPU+GPU pipeline you have the CPU, that load something in ram, then for the GPU to process it you need to move it from RAM to VRAM through a slow PCI connect then when the GPU is done you have to move it again to RAM for the CPU to handle it again. These copies cost bandwidth, latency, energy, extra memory and low level programing complexity. In the Apple unified architecture both de CPU and GPU use the same memory and avoid copying data.
You may say "ah but don't you can handle shared memory with just the DDR controller?" and that's what Intel integrated graphics do but on top of that Apple: build the CPU and GPU in the same SOC, gives the SOC a massive memory subsystem and a large system level cache.
Nvidia can do the same. Eg the Grace Hopper has the Grace CPU - nvlink C2C - Hopper/Blackwell GPU. But the trade-off is modularity. You can mix processors, ram, Nvidia GPU and all parts must work at their best capacity, but isn't even close to the fine tuning of an Apple system.
https://pages.cs.wisc.edu/~markhill/restricted/siggraph08_la...
GPUs are designed to process a large number of calculations at once (so they can process triangles in 3D graphics). This makes them good at ML applications as they can process many of the matrix calculations at once. A 4090 has 16,384 CUDA cores (general compute ALUs) and 512 Tensor cores (dedicated matrix compute ALUs); a 5090 has 21,760 CUDA and 680 Tensor cores.
The other issue when training models (and running larger models) is the amount of VRAM (or RAM for CPUs) available. GPUs are limited in this aspect, whereas CPUs can have a lot higher memory. This affects things like batch size and the size of model that can be trained or fine-tuned.
https://unsloth.ai has guides for how to fine-tune existing models like Qwen 3.8 27B, memory requirements, etc.
https://medium.com/@kailaspsudheer/the-transformers-arithmet... has some information on training a base model. A 7B llama model is estimated at taking ~34GB memory for inference at F32, but was observed requiring 96GB memory when training (for the model weights, gradients, activations, and optimizer states).
Note: you can reduce the memory required for training by recomputing the gradients, at a cost of performance/time. You can also do other tricks like performing a QLoRA/LoRA pass on the model then merging that into the model to create a checkpoint.
I don't know what sized model you could train on 64GB/128GB RAM via a CPU.
There is plenty of matrix multiplication in SIMD, but it isn't widely explored.
Like in the existing SIMD technology?
- how exactly does apple silicon s unified gpu + ram thingy work?
- how come intel and nvidia cannot do the same?
- is there an actual difference in terms of hardware architecture or something or is it pure apple marketing hype?
In a Intel+Nvidia CPU+GPU pipeline you have the CPU, that load something in ram, then for the GPU to process it you need to move it from RAM to VRAM through a slow PCI connect then when the GPU is done you have to move it again to RAM for the CPU to handle it again. These copies cost bandwidth, latency, energy, extra memory and low level programing complexity. In the Apple unified architecture both de CPU and GPU use the same memory and avoid copying data. You may say "ah but don't you can handle shared memory with just the DDR controller?" and that's what Intel integrated graphics do but on top of that Apple: build the CPU and GPU in the same SOC, gives the SOC a massive memory subsystem and a large system level cache.
Nvidia can do the same. Eg the Grace Hopper has the Grace CPU - nvlink C2C - Hopper/Blackwell GPU. But the trade-off is modularity. You can mix processors, ram, Nvidia GPU and all parts must work at their best capacity, but isn't even close to the fine tuning of an Apple system.