Microsoft Open-Sources Agent Training
5AUG
Agents now get a gym before they get a job. Microsoft Research shipped Orchard, an open framework that trains agents in realistic environments before deployment. Training, not prompting, becomes the differentiator.
Orchard separates agent training from execution. Its Kubernetes-based environment collects training data, runs reinforcement learning rollouts, and evaluates agents at scale. Three recipes ship with it: software engineering, browser use, and personal assistants.
The numbers are real. Starting from a 30-billion-parameter Qwen model, Orchard-SWE hits 67.5 percent on SWE-bench Verified, a new open-source record for its size. Code and datasets are on GitHub and Hugging Face.
Analyst John Sviokla calls it the missing layer of the agent stack. If agents can be trained like employees, model choice matters less.