M&A

Rootly buys ThinkHive to make its AI incident agents provable

What's the deal? RootlyDealroom has a profile for this one. Try Dealroom →, an AI-native on-call and incident management platform, has acquired AI agent reliability startup ThinkHiveDealroom has a profile for this one. Try Dealroom →. The deal folds ThinkHive's evaluation technology into Rootly's push to bring reliability engineering to large language model workloads. Terms were not disclosed.

What does each company do? Rootly runs incident management for customers including NVIDIA, Replit, and Canva. ThinkHive traces every step an AI agent takes and evaluates whether it did its job — not just whether it returned a response — to catch hallucination and drift in production.

Why now? Engineering teams spent a decade learning to keep distributed systems reliable. They are now deploying AI agents into production and hitting failures the legacy playbook for deterministic systems cannot handle.

The strategic rationale: ThinkHive serves a dual purpose for Rootly. It closes a customer blind spot — agents in production create incidents that legacy monitoring cannot see — and hardens Rootly's own AI, which is used during high-stakes incident response.

ThinkHive's founders learned the problem at Instacart, where co-founder Nour Alkhatib and her team did weekly forensic work to understand why agents serving millions of customers were not improving outcomes. The failures were silent. That recurring effort led Alkhatib to leave and build ThinkHive with co-founder Abdulwahab.

What changes for customers? Rootly's agents will now run on ThinkHive's evaluation engine, using groundedness scoring, hallucination detection, and shadow testing before touching a real incident. The company says this lets its agents reason from evidence to pinpoint root cause and propose fixes earlier in the lifecycle.

"When an agent gives a wrong answer at scale, that is an incident, and most teams cannot see it yet," said JJ Tang, co-founder and chief executive officer of Rootly. "We are not going to put AI into incident response and ask anyone to trust it on faith, ourselves included."

"Using AI to judge AI is like asking the same student to mark their own exam," said Alkhatib. "Real reliability means tracing what actually happened and catching the failure a score hides."

The signal: As AI agents move into production, reliability is becoming its own category. Rootly is betting that vendors will soon need to prove their agents are grounded and have not drifted — not just ship them and ask customers to trust the output.

Read more: Associated Press

Image credit: learn_tek

Source: dealroom

More top stories