NVIDIA
Desktop card
Core catalog
- GPU memory
24 GB GDDR6X
the workbench — the whole model must fit here
- Memory bandwidth
1,008 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
450 W
about 0.38 of one home's constant draw; one day (10.8 kWh) drains about 0.12 EV battery
- Form factor
Desktop PCIe card (fits a normal PC)
~$1,999
TODO: verify
In plain words: The classic enthusiast card — runs models up to ~20B parameters entirely on one card.
NVIDIA
Desktop card
Core catalog
- GPU memory
32 GB GDDR7
the workbench — the whole model must fit here
- Memory bandwidth
1,792 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
575 W
about 0.48 of one home's constant draw; one day (13.8 kWh) drains about 0.15 EV battery
- Form factor
Desktop PCIe card (fits a normal PC)
~$2,999 market price; MSRP $1,999
In plain words: The fastest thing you can put in a normal PC — 32 GB fits ~27B-parameter models with room to spare.
NVIDIA
Workstation card
Core catalog
- GPU memory
48 GB GDDR6
the workbench — the whole model must fit here
- Memory bandwidth
960 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
300 W
about 0.25 of one home's constant draw; one day (7.2 kWh) drains about 0.08 EV battery
- Form factor
Workstation PCIe card (pro tower, runs cool and quiet)
~$6,800
In plain words: Double the memory of a gaming card at lower power — for professionals who work at a desk, not in a datacenter.
NVIDIA
Server card
Core catalog
- GPU memory
48 GB GDDR6
the workbench — the whole model must fit here
- Memory bandwidth
864 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
350 W
about 0.29 of one home's constant draw; one day (8.4 kWh) drains about 0.09 EV battery
- Form factor
Server PCIe card (rack server, passive cooling)
~$11,000
In plain words: A 48 GB workhorse built for racks — the budget way to serve mid-size models from a server room.
NVIDIA
Datacenter GPU
Core catalog
- GPU memory
141 GB HBM3e
the workbench — the whole model must fit here
- Memory bandwidth
4,800 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
700 W
about 0.58 of one home's constant draw; one day (16.8 kWh) drains about 0.19 EV battery
- Form factor
SXM module (mounts on a GPU server board, 4–8 per server)
~$35,000
In plain words: 141 GB on one chip — the sweet spot for running 100B-class models without any clustering tricks.
NVIDIA
Datacenter GPU
Core catalog
- GPU memory
192 GB HBM3e
the workbench — the whole model must fit here
- Memory bandwidth
8,000 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
1,000 W
about 0.83 of one home's constant draw; one day (24 kWh) drains about 0.27 EV battery
- Form factor
SXM module (mounts on a GPU server board, 8 per server)
~$50,000
In plain words: NVIDIA's Blackwell flagship chip — 192 GB and enormous bandwidth for the biggest single-chip jobs.
NVIDIA
Full server (8× B200)
Core catalog
- GPU memory
1,536 GB HBM3e (8× B200, pooled over NVLink)
the workbench — the whole model must fit here
- Memory bandwidth
64,000 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
14,300 W
about 11.9 homes; one day (343 kWh) drains about 3.81 EV batteries
- Form factor
Complete 10U server — 8 GPUs acting as one machine
~$515,000
In plain words: One box, 1.5 terabytes of GPU memory — runs 600B–1,000B models with no clustering required.
NVIDIA
Full rack
Core catalog
- GPU memory
20,000 GB HBM3e class (72 GPUs, pooled over NVLink)
the workbench — the whole model must fit here
- Memory bandwidth
576,000 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
120,000 W
about 100 homes; one day (2,880 kWh) drains about 32 EV batteries
- Form factor
Complete datacenter rack — 72 GPUs wired as one giant machine
~$3,000,000
In plain words: A whole rack that behaves like one computer — for serving frontier-scale models to thousands of users at once.
NVIDIA
Datacenter GPU (previous gen)
Wider market
- GPU memory
80 GB HBM2e
the workbench — the whole model must fit here
- Memory bandwidth
2,039 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
400 W
about 0.33 of one home's constant draw; one day (9.6 kWh) drains about 0.11 EV battery
- Form factor
SXM module / PCIe card
~$17,000
In plain words: The chip that trained the first ChatGPT era — still a capable, cheaper workhorse today.
NVIDIA
Datacenter GPU
Wider market
- GPU memory
80 GB HBM3
the workbench — the whole model must fit here
- Memory bandwidth
3,350 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
700 W
about 0.58 of one home's constant draw; one day (16.8 kWh) drains about 0.19 EV battery
- Form factor
SXM module (4–8 per server)
~$28,000
In plain words: The GPU of the 2023–2024 AI boom — the industry's default datacenter chip.
AMD
Datacenter GPU
Wider market
- GPU memory
192 GB HBM3
the workbench — the whole model must fit here
- Memory bandwidth
5,300 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
750 W
about 0.62 of one home's constant draw; one day (18 kWh) drains about 0.2 EV battery
- Form factor
OAM module (8 per server)
~$15,000
In plain words: AMD's answer to NVIDIA — more memory per dollar than an H100, if your software stack supports it.
AMD
Datacenter GPU
Wider market
- GPU memory
256 GB HBM3e
the workbench — the whole model must fit here
- Memory bandwidth
6,000 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
1,000 W
about 0.83 of one home's constant draw; one day (24 kWh) drains about 0.27 EV battery
- Form factor
OAM module (8 per server)
~$25,000
In plain words: The most GPU memory on any single chip here — 256 GB for models that won't fit anywhere else.
Intel
Datacenter accelerator
Wider market
- GPU memory
128 GB HBM2e
the workbench — the whole model must fit here
- Memory bandwidth
3,700 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
900 W
about 0.75 of one home's constant draw; one day (21.6 kWh) drains about 0.24 EV battery
- Form factor
OAM module (8 per server)
~$16,000
In plain words: Intel's AI accelerator — priced to undercut NVIDIA, with built-in networking for clusters.
AMD
Workstation card
Wider market
- GPU memory
48 GB GDDR6
the workbench — the whole model must fit here
- Memory bandwidth
864 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
295 W
about 0.25 of one home's constant draw; one day (7.08 kWh) drains about 0.08 EV battery
- Form factor
Workstation PCIe card
~$3,500
In plain words: 48 GB of workstation memory at half the NVIDIA price — the value pick for desk-side AI work.
Tenstorrent
Accelerator card
Wider market
- GPU memory
24 GB GDDR6
the workbench — the whole model must fit here
- Memory bandwidth
576 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
300 W
about 0.25 of one home's constant draw; one day (7.2 kWh) drains about 0.08 EV battery
- Form factor
Desktop/server PCIe card
~$1,400
In plain words: The open-hardware challenger from Jim Keller's team — cheap, hackable, RISC-V based.
NVIDIA
Desktop card
Wider market
- GPU memory
16 GB GDDR7
the workbench — the whole model must fit here
- Memory bandwidth
960 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
360 W
about 0.3 of one home's constant draw; one day (8.64 kWh) drains about 0.1 EV battery
- Form factor
Desktop PCIe card (fits a normal PC)
~$1,199 (MSRP $999)
TODO: verify
In plain words: The affordable Blackwell gaming card — 16 GB runs ~13B-parameter models; the easy way into local AI.
NVIDIA
Workstation card
Wider market
- GPU memory
96 GB GDDR7 ECC
the workbench — the whole model must fit here
- Memory bandwidth
1,792 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
600 W
about 0.5 of one home's constant draw; one day (14.4 kWh) drains about 0.16 EV battery
- Form factor
Workstation PCIe card (pro tower)
~$8,500
TODO: verify
In plain words: 96 GB in a tower under your desk — runs 70B-class models locally without a server room.
NVIDIA
Desktop AI computer
Wider market
- GPU memory
128 GB LPDDR5X (unified)
the workbench — the whole model must fit here
- Memory bandwidth
273 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
200 W
about 0.17 of one home's constant draw; one day (4.8 kWh) drains about 0.05 EV battery
- Form factor
Complete mini computer (book-sized, plugs into a wall socket)
$3,999
TODO: verify
In plain words: A Grace Blackwell AI computer the size of a book — 128 GB of unified memory fits 100B-class models, though its memory is far slower than a real GPU's: for building and experimenting, not serving users.
NVIDIA
Server card
Wider market
- GPU memory
24 GB GDDR6
the workbench — the whole model must fit here
- Memory bandwidth
300 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
72 W
about 0.06 of one home's constant draw; one day (1.73 kWh) drains about 0.02 EV battery
- Form factor
Single-slot server PCIe card (no power cable needed)
~$2,500
TODO: verify
In plain words: One slot and just 72 watts — the quiet little server card for light AI work everywhere.
Intel
Datacenter accelerator (previous gen)
Wider market
- GPU memory
96 GB HBM2E
the workbench — the whole model must fit here
- Memory bandwidth
2,450 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
600 W
about 0.5 of one home's constant draw; one day (14.4 kWh) drains about 0.16 EV battery
- Form factor
OAM module (8 per server)
~$10,000
TODO: verify
In plain words: Intel's previous-gen accelerator — 96 GB of fast HBM memory at a markdown price.
Intel
Workstation card
Wider market
- GPU memory
24 GB GDDR6
the workbench — the whole model must fit here
- Memory bandwidth
456 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
200 W
about 0.17 of one home's constant draw; one day (4.8 kWh) drains about 0.05 EV battery
- Form factor
Workstation PCIe card
~$500
TODO: verify
In plain words: The budget 24 GB card — the cheapest ticket to running ~13B models at your desk.
Intel
Datacenter GPU
Wider market
- GPU memory
128 GB HBM2e
the workbench — the whole model must fit here
- Memory bandwidth
3,277 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
600 W
about 0.5 of one home's constant draw; one day (14.4 kWh) drains about 0.16 EV battery
- Form factor
OAM module / PCIe card
~$12,000
TODO: verify
In plain words: Intel's HBM flagship — 128 GB and serious bandwidth, if your software stack runs on Intel.
Tenstorrent
Accelerator card
Wider market
- GPU memory
12 GB GDDR6
the workbench — the whole model must fit here
- Memory bandwidth
288 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
160 W
about 0.13 of one home's constant draw; one day (3.84 kWh) drains about 0.04 EV battery
- Form factor
Desktop/server PCIe card
~$999
TODO: verify
In plain words: The entry ticket to Tenstorrent's open hardware — a starter card for learning the stack, not for big models.
Tenstorrent
Accelerator card
Wider market
- GPU memory
32 GB GDDR6
the workbench — the whole model must fit here
- Memory bandwidth
512 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
300 W
about 0.25 of one home's constant draw; one day (7.2 kWh) drains about 0.08 EV battery
- Form factor
Desktop/server PCIe card
~$1,399
TODO: verify
In plain words: Tenstorrent's newest generation — 32 GB for $1,399, more memory per dollar than any big-brand card here.
Tenstorrent
Desktop AI workstation
Wider market
- GPU memory
128 GB GDDR6 (4× Blackhole, networked)
the workbench — the whole model must fit here
- Memory bandwidth
2,048 GB/s
TODO: verify
how fast it re-reads the model — every word
- Power draw
1,400 W
about 1.17 home; one day (33.6 kWh) drains about 0.37 EV battery
- Form factor
Complete liquid-cooled desktop tower (4 accelerators inside)
~$12,000
TODO: verify
In plain words: Jim Keller's quiet desktop supercomputer — 128 GB across four open-hardware cards, no server room needed.