High-efficiency inference with Positron Atlas

The Positron Atlas accelerator delivers 280 tokens per second per user for Llama 3.1 8B while using only 33% of the power required by competing solutions. This photonic system addresses memory bandwidth bottlenecks and offers a sustainable alternative…

High-efficiency inference with Positron Atlas

The Positron Atlas accelerator system delivers 280 tokens per second per user for Llama 3.1 8B within a 2000W power envelope, outperforming Nvidia’s H200 in efficiency while using only 33% of the power required by competing solutions. This hardware solves the primary bottleneck for modern AI, which involves managing power consumption and memory bandwidth during inference. Positron AI uses eight Archer ASIC accelerators to achieve this throughput. The system targets organizations that prioritize sustainability and lower operational expenses.

Accelerator Throughput (Llama 3.1 8B) Power Consumption Memory Capacity
Positron Atlas 280 tokens/sec per user 2000W envelope Not specified for Atlas
Cerebras CS-3 Up to 2,500 tokens/sec ~27 kW 44GB On-die SRAM
Groq LPU 293.6 – 750 tokens/sec ~375W per chip 230MB SRAM per chip

Memory bandwidth and the Asimov roadmap

Modern AI workloads encounter significant limitations because they require immense memory bandwidth and capacity. This constraint makes the movement of model weights from memory to compute the most expensive part of the inference process. Positron AI addresses this via a memory-first architecture. The company plans to release the Asimov custom silicon to solve these scaling issues. The Asimov chip will support 2 terabytes of memory per accelerator. A Titan system will support 8 terabytes of memory. These levels will maintain realized memory bandwidth similar to Nvidia’s next-generation Rubin GPU.

The Asimov chip targets a production date in early 2027. It follows a design process that began 16 months after the company’s June Series A financing. This next-generation silicon will deliver 5x more tokens per watt in core workloads when compared to Nvidia’s upcoming Rubin GPU. The Asimov chip will also provide over 2304 GB of RAM per device. This capacity exceeds the 384 GB offered by Rubin. High-speed memory capacity per chip will provide an order of magnitude more memory than incumbent or upstart silicon providers. This capacity helps with video, trading, and multi-trillion parameter models.

You know that power consumption dictates your total cost of ownership.

Photonic processing versus electronic systems

Photonic computing represents a shift in computational architecture. Traditional electronic systems rely on the movement of electrons through semiconductor materials. Photonic systems use photons as information carriers. Light travels at the speed of light in a medium and maintains multiple wavelengths simultaneously. These systems also exhibit minimal crosstalk between parallel channels. Photonic computing provides high bandwidth, low latency, and superior energy efficiency.

The fundamental principle of photonic computing uses light-based computation to overcome the limits of transistor-based scaling. Silicon photonics platforms use CMOS-compatible manufacturing to produce photonic circuits at scale. These chips use Mach-Zehnder interferometers (MZI) to perform linear matrix multiplication. This operation forms the backbone of AI algorithms. Optical signals undergo transmission, multiplication, and interference directly on-chip using modulators, waveguides, and detectors. This process reduces latency and increases computational throughput.

One photonic processor performs 65.5 trillion Adaptive Block Floating-Point (ABFP) 16-bit operations per second. This processor consumes 78 watts of electrical power and 1.6 watts of optical power. The system integrates six chips in a single package. It uses high-speed interconnects between vertically aligned photonic tensor cores and control dies. This integration reaches the highest level achieved in photonic processing.

Precision remains a difficulty for analog processors. Standard fixed-point numbers struggle with the wide range of values in deep neural networks. The ABFP format solves this by grouping numbers into blocks. The system assigns a single, shared exponent to the entire block based on the largest absolute value. This adaptive scaling reduces quantization errors compared to simple fixed-point schemes. The system also uses analog gain control to boost lower-order bits before the signal hits the analog-to-digital converter (ADC). This technique increases precision without a more power-hungry ADC.

Comparing the inference landscape

The market for inference hardware includes several specialized architectures. Groq uses a Language Processing Unit (LPU) with 230MB of on-chip SRAM. One Groq LPU draws roughly 375W. Groq assembles hundreds of chips into an interconnected fabric to serve large models. This approach trades chip count for SRAM capacity. Cerebras uses the Wafer Scale Engine (WSE), which contains 900,000 AI cores and 44GB of on-die SRAM. The WSE-3 system draws roughly 27kW and delivers 125 PFLOPS.

Cerebras provides higher throughput than Groq. On the Llama 3.3 70B model, Cerebras documentation claims 2,100 to 2,500 tokens/sec. Independent measurements by Artificial Analysis put Groq at 293.6 tokens/sec on that same model. This creates a wide gap in speed. However, Groq provides lower costs per million tokens. For the GPT-OSS 120B model, Groq costs $0.15 per million input tokens and $0.60 per million output tokens. Cerebras costs $0.35 per million input tokens and $0.75 per million output tokens.

For the Llama 3.3 70B model, Groq costs $0.59 per million input tokens and $0.79 per million output tokens. Cerebras costs $0.85 per million input tokens and $1.20 per million output tokens. Cerebras claims a 6x price-performance advantage because of its throughput. This advantage depends on whether the user pays by wall-clock time or by token volume.

Positron AI remains a newer entrant with a less extensive ecosystem compared to established providers. The company is building an ecosystem with partners like Arm and Supermicro. It also works with industry leaders to provide infrastructure.

Edge deployment and real-time requirements

Edge AI applications require ultra-low latency for real-time decisions. This is true for autonomous vehicles, robotics, and medical diagnostics. Traditional electronic processors face limits in power consumption and heat dissipation in these environments. Photonic computing meets these needs.

Quantum Computing Inc. offers the NeuraWave platform. NeuraWave is a photonic reservoir computing platform. It debuted at SC25 and is now deployment-ready. The platform uses hybrid photonic-digital computing. This design enables real-time AI inference with ultra-low latency. NeuraWave has the form factor of a standard server PCIe plug-in card. This form factor brings photonic computing to AI at the edge. It processes data with light instead of electrons to provide real-time analysis.

The market for edge AI grows due to IoT devices and autonomous systems. Healthcare demand comes from medical imaging and wearable monitors. Photonic processors handle intensive matrix operations for medical AI without the thermal management challenges of semiconductors. Autonomous vehicle systems require sensor fusion and object recognition. The parallelism of photonic systems matches these demands. Industrial automation uses edge AI for predictive maintenance and quality control. Photonic systems provide electromagnetic interference immunity for factory floors.

Deploying sovereign AI at the edge

Organizations often need to balance performance with data residency. Positron AI partnered with i3D.net to deliver EU sovereign AI inference as a service. i3D.net provides the infrastructure and the EU footprint for enterprise customers. This partnership allows organizations to use production-grade AI inference without sacrificing data residency.

The Positron Atlas system is designed for rapid deployment. It is an inference system built for scaling. It is also a fully American-fabricated silicon and system. This allows for fast production and a dependable supply. This capability helps customers who need capacity quickly.

How will the upcoming Asimov chip change the pricing landscape for multi-trillion parameter models?

Implementing photonic workflows

Users can access these new technologies through various channels. Cerebras is available on-premises or via clouds like Meta, Vercel, Hugging Face, and OpenRouter. Groq’s LPU is available through GroqCloud and Hugging Face.

For users working in Python or R, integration involves minimal changes. Most platforms provide OpenAI-compatible APIs. Developers can switch endpoints or API keys to start testing. For those using Positron’s ecosystem, the company works with Arm to provide a broad ecosystem.

If you work with Databricks, you can connect to clusters using the databricks-connect library for Python. The sparklyr package works for R users. You must install the posit-sdk package to use Connect Viewer or service account OAuth authentication patterns.

Users can also perform local inference on consumer hardware. The GGML team, which powers llama.cpp, joined Hugging Face to support local model inference. This allows models to run directly on hardware that you control. This provides control over the infrastructure and the computation.

Specialized applications for the edge

The specialized needs of the defense and aerospace sectors drive demand for photonic edge AI. Radar signal processing and satellite communications require low size, weight, and power. Photonic systems provide radiation hardness and electromagnetic compatibility.

The industrial sector benefits from the ability to run sophisticated AI models on the factory floor. Photonic computing provides the computational density for this. This removes the dependence on cloud connectivity for immediate responses to anomalies.

Medical diagnostics also benefit from real-time optical processing. AI-assisted imaging like photoacoustics and microscopy requires ultrafast feature extraction. Photonic accelerators can perform these tasks with high efficiency.

Organizations prioritizing energy efficiency should look at the Atlas system. It delivers high throughput while using only 33% of the power of competing solutions. This reduction in energy consumption lowers the total cost of ownership. It also makes AI more sustainable for large-scale operations.

airtrain.ai
airtrain.ai

The airtrain.ai newsroom covers AI research, models and the tools built on them.

More on this topic

Stay ahead of AI

Get the week's most important AI stories delivered to your inbox every Monday.

No spam. Unsubscribe anytime.

More Stories