Sourcery

Inside Turing's RL Farms for the Agentic Age

Key points

Key takeaways from Sourcery interview with Turing CEO Jonathan Siddharth (October ​2026):

From tests to real work. The data landscape in 2026 has pivoted from helping AI pass exams — SATs, the bar, gold medals at the maths Olympiad — to helping AI master real work, with the frontier now defined by simulated reinforcement-learning environments, engineered to replicate the real world as closely as possible using prompts, verifiers and seed data from real experts.

Enterprise feedback loop. Turing doesn't just supply frontier labs with training data for coding and knowledge work: an entire division deploys agentic systems into the enterprise, and seeing how real professionals define real verifiers feeds directly back into building realistic simulated environments — a "five-dimensional matrix" of every workflow in every role in every function in every company type in every sector.

Long-horizon autonomy. Today's agents work reliably for about two days at a stretch on tasks like coding; the path ahead is agents working autonomously for weeks, months and eventually years — spinning up swarms of sub-agents, interviewing other AIs and humans,and incorporating feedback like a human colleague.

Calibrate the reward. RL environments must be tuned to the agent's sophistication: agents should succeed 20–40% of the time — too easy or too hard teaches nothing — so that the steps that earned the reward get reinforced in the neural network.

Emergent behaviour at scale. Trillion-parameter frontier models can display unanticipated emergent behaviour, as when in the Hugging Face/OpenAI incident agents passed each other messages, invented middle management and cooperated to hack — so evals and containment are essential, and scaling alone may surface behaviours that feel alien to us.

Alignment versus raw safety. Chief alignment officers and heads of safety run domain-specific evals — keeping models from reward hacking,and refusing dangerous requests across CBRN, cyber, bio, radioactive and nuclear domains — with frontier labs, in Siddharth's view, taking safety very seriously.

Cyber arms race, asymmetric bio risk. Frontier models are superhuman at discovering and patching vulnerabilities, so cybersecurity remains an active arms race where containment is an engineering problem —"we figured out how to make jet engines safe" — whereas biorisk is trickier: an engineered virus with COVID-like spread (a high K factor)and much higher fatality could outpace vaccine production and deployment.

Upleveling problems, upskilling people. AI's biggest superpower, per Alan Eustace, is raising the ceiling of problems humans can solve — from interviewing candidates on dynamic programming to asking them to "replicate Amazon" — making the skills that matter asking the right questions and verifying outputs, white frontier models help cure diseases, discover new materials and ultimately "transcend".

Read more: Sourcery

More top stories