OneTriangle launches with YC S26
Key points
What’s the deal? OneTriangle, previously known as TrustAI, is an AI infrastructure company from Y Combinator’s Summer 2026 batch. It develops inference optimisation technology based on transferring Key-Value (KV) cache state between models. A smaller model processes the initial context-heavy prefill step, then a larger model uses the transferred cache for decoding. The goal is to deliver the quality of a larger model while reducing the compute required for long inputs and improving time to first token.
Product and research: OneTriangle’s website presents the approach as a serving layer for open-weight models rather than a change to a customer’s application interface. Its research programme covers learned cache mapping, model alignment, partial recomputation and the effect of serving topology under concurrent load. In an August 2026 study, the company reported that a learned Minitron 4B-to-Llama 3.1 8B handoff made an 8K target time to first token 7.9 times faster than native prefill while improving held-out cache quality over direct reuse. Other research notes describe 91.2% retention of chance-normalised task quality in an earlier Qwen3 1.7B-to-4B experiment and continued testing across sibling model families.
Managed inference: OneTriangle also offers serviced open-weight models, handling the GPUs and serving infrastructure so customers can change one line of configuration. Its current model catalogue includes DeepSeek V4 Flash and Qwen3.6 27B, with usage-based input and output pricing. The company is led by co-founder and CEO Hannah Chung and co-founder and CTO Medhav; the founding team has backgrounds including MIT, Google DeepMind, Jane Street, SpaceX and MIT CSAIL. OneTriangle has a verified $125,000 micro-seed round from Y Combinator dated June 2026.
Read more: OneTriangle · OneTriangle research · OneTriangle models and pricing · Y Combinator video