The AI Memory Wall: How Photonics Could Reshape the Future of AI Chips

Prefer to listen instead? Here’s the podcast version of this article.

Artificial intelligence has spent the past few years in a relentless race for more computing power. Bigger GPUs. Bigger clusters. Bigger data centers. Bigger electricity bills.

 

But the next major AI performance breakthrough may come from something less glamorous: helping processors get data from memory faster.

 

San Francisco semiconductor startup Volantis has raised $88 million in Series A funding to tackle exactly that problem. Rather than attempting to build yet another GPU to compete directly with NVIDIA or AMD, Volantis is developing photonic technology that uses light to connect AI computing chips with much larger pools of memory.

 

According to [Reuters] the company believes its optical approach could allow a GPU to communicate with as many as 220 memory chips, dramatically expanding the amount of memory that can sit within high-speed reach of an AI processor.

 

That could matter enormously as AI moves from model training toward always-on inference, coding agents, long-context models and autonomous enterprise workflows.

 

In other words, AI’s next speed limit may not be how quickly a GPU can calculate. It may be how quickly the GPU can get the data it needs.

 

 

The $88 Million Bet on the AI “Memory Wall”

Volantis says its Series A was led by Lachy Groom and Abstract Ventures, with participation from John Doerr, VXI Capital, Triatomic and Susa Ventures, alongside angel investors including Dwarkesh Patel, Naveen Rao and Sholto Douglas. The round takes the company’s total funding to approximately $97 million.

 

The company explains the strategy in its own post [Volantis] Rather than treating memory merely as storage sitting beside a processor, Volantis wants to rethink the physical connection between memory and AI compute.

 

That distinction is important.

 

Modern AI accelerators can perform an astonishing number of mathematical operations every second. But large language models contain enormous volumes of parameters and intermediate data that constantly need to move between memory and the processor.

 

When the processor has to wait for that data, expensive compute capacity can sit underutilized.

 

Welcome to the memory wall.

 

 

Why AI Chips Need So Much Memory Bandwidth

Think of a GPU as an extremely fast chef.

 

The chef might be capable of preparing hundreds of meals every hour, but that speed is useless if ingredients arrive through a tiny kitchen window one tomato at a time.

 

AI processors face a similar problem.

 

Model weights, key-value caches, context and other data have to move continually between memory and computation. As models become larger and context windows expand, the amount of information traveling through the system grows substantially.

 

That’s why high-bandwidth memory, or HBM, has become critical to modern AI accelerators.

 

Micron offers an excellent technical explanation in its [blog on how memory os becoming the backbone of the AI era] Its broader argument is worth paying attention to: AI system performance is increasingly determined not only by computation but by bandwidth, latency and data movement.

 

The established GPU industry is already responding aggressively.

 

For example, NVIDIA’s next-generation Vera Rubin platform incorporates HBM4 and is designed for up to 288 GB of memory and as much as 22 TB/s of aggregate memory bandwidth per GPU, according to the company’s technical materials. NVIDIA’s illustrates how central memory bandwidth has become to accelerator design.

 

Volantis is proposing a different way to push those limits.

 

 

Instead of Copper, Volantis Wants to Move AI Data With Light

Traditional chip packages predominantly move information through electrical connections.

 

The catch is that extremely high-speed electrical connections become increasingly difficult as distance, bandwidth, energy consumption and signal integrity requirements rise.

 

Volantis wants to replace some of those short electrical connections with optical links.

 

Its technology uses tiny lasers known as vertical-cavity surface-emitting lasers, or VCSELs, to transmit information through light rather than conventional electrical wiring.

 

VCSELs are not entirely experimental technology. They have already been manufactured at massive scale for applications including smartphone sensing systems. Reuters notes that related VCSEL technology has been used in Apple’s Face ID ecosystem, providing Volantis with a comparatively mature manufacturing foundation rather than requiring an entirely new laser supply chain.

 

The clever bit is where Volantis wants to put the technology: inside the AI accelerator’s memory architecture.

 

Its technical materials say electrical interposer connections typically reach only a few millimeters, while Volantis’s optical waveguides are designed to extend beyond 200 millimeters. The company says this additional reach could make it possible to connect more than 220 memory chiplets into one uniform-latency memory pool.

 

That is a radically different geometry for AI hardware.

 

Instead of cramming a small number of extremely expensive memory stacks as close as physically possible to a GPU, photonics could potentially create a much larger neighborhood of memory surrounding the compute engine.

 

 

What Makes Volantis Different?

Volantis is not simply talking about putting optical cables between racks of GPUs.

 

Its pitch goes much closer to the processor.

 

The company says its optical fabric uses integrated micro-VCSEL arrays and extremely small optical waveguides rather than conventional fiber and external lasers. It lists target characteristics including 24 Gb/s per lane, sub-five-nanosecond latency, tens of thousands of lanes and more than 200 TB/s of aggregate bandwidth.

 

Volantis also says its links consume less than one picojoule per bit and have demonstrated wafer-scale bit-error rates below 1e-12. These figures come from Volantis itself and should therefore be understood as company-reported technical results and design characteristics rather than independent third-party benchmarks.

 

That distinction matters because semiconductor engineering has a habit of turning beautiful diagrams into very difficult manufacturing problems.

 

Packaging yields, thermal behavior, laser reliability, software compatibility, manufacturing cost and large-scale system integration will all matter if Volantis wants its technology to move from promising silicon into production data centers.

 

 

The A-1 System Targets Massive AI Models

Volantis calls its first planned system A-1.

 

The company’s announcement says A-1 is being designed for extremely large AI inference workloads, with ambitions extending into multi-trillion-parameter models and extremely high per-user token throughput. The company’s website currently promotes support for models above 10 trillion parameters and speeds of up to 10,000 tokens per second per user, while its formal funding announcement distributed through PR Newswire describes A-1 as being designed for models exceeding 20 trillion parameters.

 

Those are targets, not results that prospective customers should assume have already been demonstrated in a commercial production system.

 

That caveat aside, the goal reveals where Volantis sees the AI market heading.

 

The company is focusing heavily on inference rather than simply training.

 

That is important because the AI economy is increasingly about running models repeatedly for millions of users and software agents—not merely training a frontier model every few months.

 

 

Why Faster Memory Could Supercharge AI Agents

Imagine an AI coding agent asked to analyze a large software repository.

 

The agent may need to ingest hundreds of files, reason across multiple dependencies, generate code, run tests, inspect results and revise its solution.

 

That workflow isn’t one simple model response. It is a long chain of inference operations.

 

When every step waits on model execution, latency stacks up.

 

Volantis argues that dramatically increasing memory bandwidth and capacity could make these workflows far more interactive. Its own funding announcement uses coding agents as an example of applications that could potentially move from lengthy workflows toward near-real-time operation if model inference becomes dramatically faster.

 

 

Energy Efficiency Could Become Just as Important as Speed

There is also a sustainability angle.

 

Modern AI systems spend considerable energy not just performing mathematical operations but moving information among processors, memory and networking hardware.

 

As infrastructure expands, optimizing data movement could become a meaningful component of reducing the energy required per AI query.

 

Micron’s research has similarly emphasized memory as an important frontier for AI energy efficiency, noting that memory, storage and the interconnects surrounding them can represent a substantial portion of system power depending on the workload and architecture.

 

Optical connections are therefore interesting for reasons beyond benchmark bragging rights.

 

If photonics can move substantially more information using less energy per bit, improved interconnects could help data centers generate more useful AI work from constrained electrical power.

 

That matters because electricity availability is rapidly becoming an infrastructure constraint of its own.

 

 

The Challenges Volantis Still Has to Prove

An $88 million funding round is a major vote of confidence. It is not the same thing as volume production.

 

Volantis now has to demonstrate that its architecture works reliably and economically at scale.

 

That means proving several things simultaneously: the photonic components must achieve high manufacturing yields; packaging must tolerate the thermal environment surrounding powerful AI processors; the memory system must deliver reliable latency and bandwidth under real workloads; software needs to take advantage of the architecture; and the final system needs to produce sufficiently compelling total-cost-of-ownership improvements to justify infrastructure changes.

 

The competitive landscape won’t stand still, either.

 

NVIDIA, AMD, memory manufacturers and advanced-packaging suppliers continue improving HBM capacity, bandwidth, chiplet architectures and scale-up networking. NVIDIA’s Rubin architecture, for example, shows just how aggressively conventional HBM-based systems are evolving.

 

So Volantis does not merely have to prove that photonics works.

 

It has to prove that photonics improves quickly enough to outrun improvements in the technologies it hopes to disrupt.

 

 

What the Volantis Funding Means for the Future of AI

The most interesting thing about Volantis’s $88 million raise is not necessarily Volantis itself.

 

It is what the investment says about where the AI industry believes the next bottlenecks are emerging.

 

The first phase of the generative AI boom revolved around models.

 

The second became a race for GPUs.

 

The next phase is becoming a full-stack infrastructure battle involving memory, networking, advanced packaging, optical interconnects, cooling, energy and highly specialized silicon.

 

That changes how technology leaders should think about AI strategy.

 

The fastest model on paper means very little if infrastructure cannot feed it enough data, power it economically or operate it at acceptable latency.

 

For CIOs and enterprise technology leaders, this means AI infrastructure decisions should increasingly consider memory architecture, inference efficiency, networking and cost per token, not simply which accelerator produces the largest benchmark number.

 

For semiconductor companies, it means some of the industry’s biggest opportunities may exist between the processor and memory rather than inside the processor itself.

 

For the wider AI ecosystem, Volantis provides another signal that the next generation of innovation will involve as much physics as software.

 

 

Conclusion

Volantis’s $88 million funding round highlights a major shift in the AI hardware race: the industry is no longer focused only on building faster processors. The ability to move enormous amounts of data between AI accelerators and memory is becoming just as important as raw compute performance.

 

By using photonic interconnects to link AI chips with larger pools of memory, Volantis is taking aim at the growing “memory wall” that can limit the performance and efficiency of advanced AI systems. If the company can successfully deliver its promised gains in bandwidth, latency, energy efficiency, and memory capacity at commercial scale, its technology could help reshape how future AI infrastructure is designed.

 

There are still significant technical and manufacturing challenges ahead. Volantis will need to prove that its architecture can perform reliably in real-world data centers, compete economically with rapidly improving HBM-based systems, and integrate smoothly with the broader semiconductor ecosystem.

 

Even so, the company’s funding sends a clear signal about where AI infrastructure innovation is heading. The next breakthroughs may come not only from more powerful GPUs, but from smarter ways of connecting compute, memory, networking, and storage.

 

As AI models become larger and inference workloads become more demanding, technologies that reduce data-movement bottlenecks could become critical to lowering costs, improving performance, and making advanced AI systems more accessible.

 

Volantis is still early in that journey, but its photonics-first approach places it squarely in one of the most important emerging battles in AI hardware: making sure tomorrow’s processors can access data as quickly as they can process it.

WEBINAR

INTELLIGENT IMMERSION:

How AI Empowers AR & VR for Business

Wednesday, June 19, 2024

12:00 PM ET •  9:00 AM PT