On September 15, Lightbits Labs unveiled Inferra, a new software engine designed to address one of the most pressing bottlenecks in artificial intelligence inference: the memory wall created by expanding context windows and multiplying user sessions. The company, known for pioneering NVMe over TCP technology, is applying that expertise to help data centers and cloud providers squeeze far more computational value from existing GPU hardware.

The Problem Inferra Solves

Modern large language models store attention states, or KV caches, in GPU high-bandwidth memory (HBM) during inference. As these models process longer documents and handle more concurrent users, the KV cache balloons in size. When it exceeds available GPU memory, systems face a costly choice: either discard the cache and recompute it later, burning GPU cycles, or reject additional user sessions. Both paths waste expensive computational resources.

Inferra virtualizes GPU memory across DRAM and NVMe storage tiers, transforming the KV cache into a persistent layer that spans multiple hardware components. Rather than losing attention states when they exceed HBM capacity, the engine keeps them alive on NVMe SSDs and intelligently prefetches them back to the GPU just before the model needs them.

Performance Claims and Real-World Testing

AI inference processing memory cache
Photo by Hitesh Choudhary

Lightbits claims the engine delivers up to 16 times greater concurrent inference session density on existing GPUs. The company also reports latency improvements exceeding 100x compared to recomputing attention states from scratch. Context windows can expand to 10 million tokens, far beyond what fits in typical GPU memory pools.

These are ambitious figures, and Lightbits has already tested them in customer beta environments, including with OVH, a major European hosting provider. OVH reported substantial GPU utilization gains, enabling more scalable and cost-effective infrastructure for AI agent and retrieval-augmented generation (RAG) workloads. Solidigm, the SSD manufacturer, ran Inferra in its AI Central Lab using its D7-PS1010 drives, demonstrating how network-attached NVMe storage paired with intelligent software virtualization can overcome memory constraints for agentic AI applications.

Technical Architecture

enterprise NVMe storage systems
Photo by Marc PEZIN

Inferra sits between serving frameworks (vLLM, TensorRT, and SGLang are supported) and a standard pool of NVMe SSDs. The engine includes four core components: an intelligent prefetcher that predicts when attention states will be needed, a tiering manager that shuffles data between memory and storage tiers, a log-structured KV store optimized for flash media, and security and isolation engines.

The log-structured design is critical for practical operation. Rather than performing small random writes that would degrade NVMe endurance, the engine appends updates sequentially, minimizing wear. This approach allows standard enterprise SSDs to serve as an effective KV cache tier. For multi-tenant environments, Lightbits has engineered secure tenant isolation and encrypted data transfer between tiers, enabling hosting providers to allocate a single NVMe pool to multiple customers with consistent performance guarantees.

Companies deploying NVMe-based infrastructure can apply this software to reduce reliance on GPU expansion or costly memory upgrades.

What This Means for Buyers

For enterprises and cloud providers running inference workloads, Inferra addresses a growing pain point. As AI adoption accelerates and models grow larger, memory becomes the constraining resource. This software-defined approach lets organizations extract more value from current GPU investments without hardware changes.

The engine is currently in customer beta with no publicly announced pricing or general availability date. Lightbits is demonstrating results at industry events and offering a Pod Efficiency Analyzer tool for teams wanting to model potential GPU savings against their own cluster configurations before committing.

The fundamental shift Inferra represents is significant: moving away from treating inference as a simple extension of training-optimized systems toward purpose-built infrastructure. Lightbits argues that legacy storage and serving stacks designed for model training cannot efficiently handle the unique demands of multi-session, long-context inference. By decoupling KV cache storage from GPU HBM and adding predictive intelligence, the company is offering a way to rebalance the economics of AI deployment.

For buyers evaluating inference infrastructure or planning capacity upgrades, understanding how NVMe storage at scale integrates with orchestration software will become increasingly relevant. Inferra is one of the first purpose-built engines in this emerging category, and its performance metrics, if independently verified, could reshape how data centers prioritize memory and storage when provisioning AI clusters.