20VC

Positron's Thomas Sohmers: AI inference's real bottleneck is memory, not compute

Key points

Key takeaways from a 20VC interview with Positron AI Co-founder Thomas Sohmers (September 2026), weeks after the company announced an $875M Series C at a $5B valuation to scale its energy-efficient AI inference silicon:

Inference is a memory problem, not a compute problem. Between 2014 and 2024 single-GPU flops improved roughly 120x while memory bandwidth improved only about 17x, widening the "memory wall". Because generative inference must read model weights for every token, it is heavily memory-bound, unlike training, which parallelises easily.

KV caching is where inference economics sit. On SemiAnalysis' agentic-coding benchmark of real Claude Code sessions, roughly 96% of tokens are cached, so persistent per-user caches become an operator's biggest lever. At long context windows, single user sessions can reach ~100GB — with multi-trillion-parameter frontier models around ~5TB of weights, 50 active users' contexts can exceed the model itself.

Positron's answer is memory-first silicon. Its next generation ships roughly 8x the memory per device of Nvidia's highest-memory SKU — which Sohmers says Nvidia is actually shrinking on cost grounds — and the goal is ~5x more compute per megawatt: doing in ~100MW what would need ~500MW of Nvidia equipment.

Energy is real, but not the binding constraint. Sohmers argues the bigger limiters are economics and debt, not raw generation capacity: new data centres typically bring generation that covers their own use, and he calls the mainstream anti-data-centre backlash a scapegoat built on false premises, such as the In-N-Out water myth, while US federal land in Nevada and the West offers room for clean buildout.

He is opposed to "pacing the frontier". Sohmers is sceptical of pause-driven regulation — he fears concentrating AI capability in a few companies and governments — and notes labs hardly need a pause for profitability: Anthropic is reported to run around 80 points of gross margin on its API business.

Token prices collapsed while token value exploded. The Silicon Data token price index fell from ~$60 per million tokens five years ago to below $1 per million this month, but Sohmers argues capability per token is up 100–1000x, and pricing may eventually shift from per-token toward per-useful-result.

Frontier models keep scaling — and local models feed them. Roughly 80–85% of tokens today come from the top four model companies; he expects frontier sizes to keep growing since the scaling laws are observed, not proven, to hold — while on-device models will fire more, not fewer, requests at cloud leaders like OpenAI. He credits Chinese labs, forced by export controls, with the sharpest recent inference innovations (DeepSeek's multi-head latent attention and gated delta nets), noting there is no free lunch in capability tradeoffs.

Read more: 20VC on YouTube

More top stories