RISC-V data movement and Tensix cores
Tenstorrent provides a legitimate RISC-V alternative to NVIDIA for developers requiring full-stack auditability. The Tensix core is comprised of five small RISC-V cores, a matrix engine, and a vector engine. These five small cores handle ingestion of operands, unpacking, compute, packing, and output back to the network-on-chip. Each Tensix core is equipped with 1.5 MB of local L1 SRAM. The architecture is built on a 2D mesh of Tensix cores. Unlike NVIDIA’s SIMT model, Tenstorrent requires programmers to schedule data movement via DMA operations. This explicit movement replaces the hardware prefetcher. The Blackhole architecture is built with 16 big RISC-V cores to manage data orchestration on-die. This design eliminates the host CPU overhead that slowed Wormhole during small-batch workloads. The matrix engine is an 8×8 BF16 tile multiply-accumulate unit, and the vector unit is used for element-wise operations such as activations and normalizations. The 6nm manufacturing process for Blackhole is used to increase density and network-on-chip speeds.
Hardware products and performance specs
The product lineup is designed for various deployment scales. The Wormhole n150 PCIe card is a single-chip solution with 72 Tensix cores, 108 MB of SRAM, and 12 GB of GDDR6 memory at 160W. The dual-chip Wormhole n300 board is a 300W system with 128 Tensix cores, 192 MB of SRAM, and 24 GB of GDDR6 memory. The Blackhole p100 card is a $999 option with 28 GB of GDDR6 and 120 Tensix cores. The Blackhole p150 developer card, which is equipped with 32 GB of GDDR6 and 120 Tensix cores, is priced at $1,399 and is available in active, passive, and liquid-cooled variants for diverse thermal and mounting requirements. A February 2026 firmware update (v19.5.0) reduced the p150 core count from 140 to 120 and decreased SRAM from 210 MB to 180 MB. This change is a result of silicon yield management on the 6nm process. The TT-Quietbox 2 desktop workstation is a liquid-cooled system that runs four Blackhole processors to deliver 2,654 TFLOPS of BlockFP8 compute. This $9,999 system is capable of running models up to 120 billion parameters. It is equipped with 128 GB of GDDR6 and 256 GB of DDR5 memory. The Quietbox 2 is able to run Llama 3.1 70B at 476.5 tokens per second. It is also able to predict four protein structures in parallel using the Boltz-2 model in 49 seconds, whereas a modern CPU is able to take 45 minutes for the same task.
| Product | Memory | SRAM (per chip) | Peak FP8 Compute |
|---|---|---|---|
| Wormhole n150 | 12 GB GDDR6 | 108 MB | 262 TFLOPS |
| Wormhole n300 | 24 GB GDDR6 | 192 MB | 466 TFLOPS |
| Blackhole p100 | 28 GB GDDR6 | 180 MB | 745 TFLOPS |
| Blackhole p150 | 32 GB GDDR6 | 180 MB | 745 TFLOPS |
Scaling and the software stack
The TT-Forge compiler is an MLIR-based frontend that handles graph optimization, lowering, and code generation. It is able to run models from PyTorch, ONNX, TensorFlow, JAX, and PaddlePaddle. Tenstorrent scales compute using an on-chip Ethernet mesh instead of proprietary interconnects. The Blackhole Galaxy rack system is a 6U system that uses 32 chips to deliver 23 PFLOPS of FP8 compute. This system uses 56 x 800G Ethernet ports to enable 11.2 TB/s of scale-out bandwidth. The Galaxy configuration is equipped with 6.2 GB of on-chip SRAM and 1 TB of DRAM. A 16-unit Galaxy cluster running DeepSeek R1 671B is able to deliver 350+ tokens per second per user and achieve a sub-4-second time-to-first-token on 100K context. This performance is achieved at a cost of $6 per million tokens. You might wonder if the software maturity can ever match CUDA. The hardware is effective for premium tokens where low latency is required for agentic workflows and video generation. The software ecosystem is still behind NVIDIA, which makes model deployment a difficult engineering project. The Wormhole Galaxy is able to reach 4,000 to 5,000 tokens per second for Llama 70B at batch 32, which is higher than the 2,500 to 3,500 tokens per second achieved by an 8x H100 SXM5 node running vLLM.




