LlamaIndex doubles down on the context layer for always-on agents
Key points
Key takeaways from theCUBE interview with LlamaIndex CEO Jerry Liu at CoreWeave Fully Connected 2026 in San Francisco (September 2026):
From framework to specialised lab. LlamaIndex launched three years ago as a popular open-source framework for building early RAG and agentic workflows; it has since narrowed its mission to the context layer and now operates as a specialised AI lab, post-training open-weight models purely for document parsing and extraction.
Millions of pages a day. LlamaIndex processes millions of pages of documents each day for finance, legal and insurance customers, with highly bursty traffic; the workload mix is roughly 75% inference and 25% training, and Liu expects inference demand to grow faster than training.
Why CoreWeave. Latency at scale and absorbing extreme traffic bursts were decisive in choosing CoreWeave as compute partner; Lukas Biewald, now SVP of AI Initiatives at CoreWeave, says how chips are networked and powered can shift performance by orders of magnitude, visible in public benchmarks, and points to close collaboration with Nvidia on data-centre design.
Always-on agents. The two see AI shifting from short Q&A tasks to long-running, always-on agents that keep context and memory; "you wouldn't hire an employee for one day", Liu argues, so persistent agents with memory unlock far more applications.
Governance before scale. Biewald pushes back on the idea of "rogue agents": they do what they are told, so enterprises need explicit privacy and scope controls, and governance is currently one of the biggest constraints on wider enterprise adoption.
Evals for every task. Teams should build evaluation benchmarks for every task, measuring alignment with human goals rather than accuracy alone; prompt engineering is moving from prescriptive instructions towards broad scaffolding with clear reward signals.
Observability heritage. Liu and Biewald first worked together three years ago on a native integration between LlamaIndex and Weights & Biases observability; after CoreWeave acquired Weights & Biases about a year and a half ago, the team launched hardware-failure visibility so researchers can see where models break, not just whether they work.
Read more: theCUBE on YouTube