Discussion about this post

User's avatar
State of Play's avatar

Podjarny's line about skills needing code's lifecycle lands harder against the production numbers: 88 percent of agent pilots never ship, and the pilots that do share almost nothing except verification infrastructure sized to match the generation side. That's his eval-versus-eyeball distinction, just measured at the fleet level instead of the file.

The harness-ownership argument tracks the same shape. Orgs actually running agents in production aren't buying more capable models, they're buying review capacity for what the models already produce, which is closer to owning your constraints layer than to a model upgrade. Cost and activity tracking are turning into a separate line item as verification catches up to generation, which is the same gap this piece is describing one level down, at the skill.

No posts

Ready for more?