How to Build Long-Horizon AI Agents — Mitch Troyanovsky, Basis
The MAD Podcast with Matt Turck
AI agents can write code for hours, but ask them to do real work in the real economy, and they break. Mitch Troyanovsky is co-founder of Basis, a unicorn AI company whose agents run autonomously for hours — sometimes days — completing complex tax returns end to end. His answer to the reliability problem: stop grading outcomes, and start supervising the process.
This is a definitive, reference-style conversation on building long-horizon AI agents. Mitch walks through the full history — from ReAct and the AutoGPT crash to reasoning models and RLVR — and explains why the industry abandoned process supervision in 2023, and why it's now coming back at a completely different scale. We go deep on behavior specs, the open standard Basis just released with Braintrust for defining and evaluating how agents behave across entire trajectories, with no ground truth required.
Along the way: why context is really runtime training data, why your documentation must be treated like a codebase, ontologies as "worlds for agents to live in," the judge-as-agent architecture, why Basis hires philosophy majors as Language Architects, deploying agents as "onboarding 300 brilliant alien employees," and Mitch's prediction for when the bitter lesson swallows the harness.
(01:09) Why Basis Engineers Whisper to Their Agents
(04:12) Accounting as Compression: an Intelligence Layer Over the Economy
(06:11) Defining Long-Horizon: When You Exceed the Context Window
(08:24) Anatomy of a Multi-Day Autonomous Trajectory
(10:19) Handoff Design: Optimizing Output for the Reviewer
(11:17) ReAct and Why Reasoning Must Regulate Its Own State
(12:33) Large Working Memory, No Long-Term Memory
(14:13) Compounding Errors: Why AutoGPT and BabyAGI Broke
(15:51) Opus 3, o1, o3: the Three Real Paradigm Shifts
(17:07) Titrating Inference Compute Across Easy and Hard Steps
(18:23) Process Reward vs. Outcome Reward: "Let's Verify Step by Step"
(20:32) RLVR and Why the METR Curve Overstates Reliability
(22:09) Verifiable at Runtime: the Real Reason Coding Won
(25:14) No Ground Truth, No Cheap Verification, No Data
(26:55) Encoding Deterministic Checks From Human Review Process
(29:18) Synthetic Data Limits: Generating Artifacts, Not Text
(33:16) 100 Evals Pass — Does It Generalize to Production?
(35:53) Primary Sources vs. Pre-Training Knowledge
(36:37) Behavior Specs: Markdown, Judges, and True/False/N.A.
(39:58) Specificity vs. Brittleness in Spec Authoring
(42:18) Context as Runtime Training Data
(44:21) Judge-as-Agent: Trajectory Maps and Sub-Agent Attribution
(46:45) The Move 37 Objection: Reliability Over Optimality
(50:02) The Magic Box Model: Building Without Weights Access
(52:41) "Nothing Paradigm-Shifting Has Changed Since o3"
(54:56) Open-Sourcing the Behavior Spec Standard With Braintrust
(01:02:54) Ontology Design: Virtual Filesystems, Graphs, Embeddings
(01:04:20) Canonical vs. Non-Canonical: Docs as Codebase
(01:06:33) Language Architects and Writing for Runtime Interpretation
(01:09:05) Deployed Intelligence: 300 Alien Employees With No Context
(01:11:10) Closing the Loop: Signal → Context, Tools, Harness
(01:12:50) Context Slop: the Mistake Most Agent Builders Make
(01:14:29) Reward Function Design and Credit Assignment Over Trajectories
(01:17:01) Will the Bitter Lesson Swallow the Harness?
(01:18:46) Business Moats vs. Technical Moats
(01:21:03) Paradigm Thinking Over Timeline ADHD