Memory layer for agentic inference

Your GPUs are
waiting on memory.

As AI shifts from chatbots to agents, a single session's state holds scarce HBM while the agent waits seconds on every tool call. NullHz is a vendor-neutral control plane that moves KV/session state across the memory hierarchy just in time — so the GPU never stalls, and you serve materially more concurrent agents on the hardware you already own.

More Concurrent agents
Zero GPU memory stalls
Any Vendor / engine
0 New hardware
Works across

The Problem

The most expensive resource in the data center sits idle.

Agents don't run flat out like a chatbot. They think, then wait — seconds at a time — on each tool call. But their session state never leaves HBM. That stranded capacity is the quiet tax on every agentic deployment.

Agents wait. GPUs idle.

A single agent waits seconds on every tool call. The GPU has nothing to do — yet that session's KV state still pins scarce HBM the entire time.

HBM is the bottleneck.

HBM is the scarcest, most expensive memory in the building. Stranded session state — not compute — is what caps how many concurrent agents you can serve.

Overprovisioning is the only lever.

To hit concurrency targets, operators overprovision GPUs — and still leave utilization on the table. In today's supply-constrained market, that capacity can't be bought at any price.

How It Works

An intelligent control plane for session state.

NullHz manages the full lifecycle of KV/session state across the memory hierarchy and across nodes — deciding where each session's state should live, and moving it just in time so the GPU never stalls waiting on memory.

NullHz control plane workload-aware placement & just-in-time movement
HBM fastest · scarcest

Active sessions that need the GPU now live here. Everything else is evicted to make room.

Host memory warm · abundant

Sessions paused mid-tool-call park here — milliseconds from being promoted back to HBM the instant they resume.

NVMe cold · cheap · deep

Long-idle sessions spill to local NVMe and across nodes, freeing HBM for live work without dropping state.

Tiered, end-to-end

Manages KV/session state across HBM, host memory, and NVMe — and across nodes — as one coherent hierarchy.

Just-in-time movement

State is pre-staged so it's back in HBM the moment an agent resumes. The GPU never blocks on a memory fetch.

Workload-aware

Decisions are driven by each session's behavior — tool-call latency, idle patterns, priority — not a fixed, dumb LRU.

Vendor-neutral

Sits above the transport layer and alongside your serving engine. No lock-in to a single GPU vendor or runtime.

Why NullHz

The workload-aware tier serving engines leave open.

NVIDIA's transport layer moves bytes between memory tiers extremely well. Serving engines schedule tokens. Neither one owns the decision of where each agent's session state should live over its whole lifecycle — and that's exactly the gap that strands HBM.

NullHz sits in that tier. It doesn't compete with the transport layer — it sits above it, making the workload-aware placement and movement decisions that turn stranded capacity back into served agents.

Request Early Access
Complements, doesn't compete

Sits above NVIDIA's transport layer rather than replacing it.

Full-lifecycle management

Owns session state from first token to eviction, across the whole hierarchy.

Capacity you can't buy

Reclaims stranded HBM to serve more agents on the GPUs you already own.

Drop-in tier

Vendor- and engine-neutral, so it fits the stack you're already running.

↑ Concurrency Serve materially more agents per GPU
↑ Utilization Turn idle HBM into useful work
↓ Spend Defer GPU buys you can't make anyway

Early Access

Reclaim your stranded GPU capacity.

We're partnering with a small group of operators running agentic workloads at scale. Tell us about your stack and we'll show you how much HBM NullHz can give back.

Hands-on evaluation with our team
Vendor- and engine-neutral — fits your stack
Direct line to the founding engineers