Tristan Handy and Jason Ganz are back after a great time at dbt Summit. In this episode, they discuss four topics: the arrival of the personal agent from Instinct, Muse, Dot, and others; a sharp shift in public opinion on AI risk; OpenAI canceling the planned release of GPT-6.1 Astra after alignment testing; and Jev, a new “decision model” from TypeSafe that outputs calibrated probabilities instead of tokens.
Please reach out at podcast@dbtlabs.com for questions, comments, and topics you want covered next time.
dbt State is now generally available. Once it’s on, dbt State checks what has changed upstream, in both code and data, and skips rebuilding the models that haven’t. That’s compute you’re not paying for, and no more waiting on upstream models you aren’t touching in dev. It works wherever you use dbt. Try dbt State with a 30-day free trial for eligible new organizations.
Listen now: Spotify · Apple Podcasts · YouTube · Amazon Music · RSS
Topic one: the era of the personal agent
Jason argues that we’ve hit a tipping point for personal agents: assistants that work with one person continuously, with access to their personal and work context. He walks through the new entrants, including Instinct (a startup that reportedly reached a $10 billion valuation), Muse from Meta Superintelligence Labs, Grokbot, and Dot from OpenAI. Tristan traces the lineage back to OpenClaw, explains what’s different now (managed harnesses and cloud compute instead of a dedicated Mac mini), and describes the personal agent he built himself with tight permissions and its own email account.
Condensed transcript
Jason Ganz: One of the things that has been on my mind for a while is when we’re going to actually enter the era of the personal agent. As soon as agents became a thing, we had the idea of a chief of staff agent that manages your communications and your calendar, an ongoing, continuous presence in your life with access to your personal context and probably your work context. There have been a couple of false starts, mostly in the realm of early adopters like myself. But over the past few weeks we’ve seen a pretty clear tipping point. People are adopting it and the major players are lining up behind it.
Tristan Handy: What I find weird is that all of this is happening at the same time. To me this really started with OpenClaw, the first real one of these that took off with very tech-forward power users. It required you to run your own dedicated hardware, which is when there was a run on Mac minis. The big differences now are that the agent harness is managed for you, which also means it’s out of your control, and that they provide cloud compute, so you don’t need to provision your own hardware.
Tristan Handy: Maybe this is my elder millennial preferences showing, but I get most things done via email. My productivity life, which is what I actually want an agent to help with, happens in email. If I can forward an email and say “pay this vendor,” that’s useful work. But none of these currently have email addresses you can set up. That plus a few other things made me build my own personal agent a couple of months ago. It runs on a Mac mini in my basement, has its own agent mail account, and I built the harness and skills specifically around permissions: I want it to handle high-criticality things, but I want to be very specific about what it can and cannot do, and where it must come back to me.
Jason Ganz: Where I think this is most useful for me is managing the astonishing amount of information inputs we have as knowledge professionals. I had a Claude Cowork automation that, before I checked Slack every morning, would look across Slack with some context about which projects I cared about and say, here are the three things that need your attention. If I open Slack cold, I’m pulled into a vortex of 70 things that have happened since yesterday. Being able to filter to the most important thing you should be working on right now is a tremendously impactful interface change, and hopefully the end of the nine-billion-unread-notifications era.
Tristan Handy: My version of that is one vibe coding session away. I want to forward every email related to my kids’ school to this agent and have it tell me what I actually need to know. The volume of communication is stunning, and most of it I don’t care about, but there’s that one line that says “make sure you help them with this particular homework.” I think this category will be dominated by the big consumer tech companies or the model labs. Because those products are built for consumers, they don’t have the control I personally want, so I’m not sure I’ll be a long-term user of any of them. It turns out it’s just not that hard to build something.
Topic two: AI becomes a top issue for voters
Tristan walks through polling summarized in Zvi Mowshowitz’s post “The AI Preference Cascade Reaches Farther”: the share of voters naming AI their single top issue has gone from zero last year to 3% in March to 8% now, and the top concern has shifted from job losses to humans losing control of AI. He shares highlights from a recent Senate subcommittee hearing on the Hugging Face incident, including a senator asking whether recursive self-improvement (RSI) should be made illegal. Jason reacts to the speed of the shift, notes how small Metr is relative to the stakes, and raises the concern that legislative timescales don’t match the timescales some lab insiders are predicting.
Condensed transcript
Tristan Handy: A year ago the general public was moderately curious about AI and generically split on optimism versus pessimism. Now this is one of the top concerns for voters. That’s not to say it’s nearly as high as cost of living, but 8% of people polled say AI is their single top issue, up from 3% in March, and it didn’t exist as a top issue in this poll last year. For most of these issues the numbers are very stable, so moving that aggressively is really notable. Between March and September the number one concern changed, from AI causing major job losses and economic disruption to AI eventually becoming powerful enough that humans lose control over it. I would not have predicted that. And 62% of voters now say AI development should slow down.
Tristan Handy: There was a Senate subcommittee hearing last week about the Hugging Face incident, and the head of Metr was being questioned. At one point a senator asked, flat out, “Should we just make RSI illegal?” Every single senator seemed to think we needed both harsher liability regimes for AI developers and new legislation, both very quickly. The room was concerned about falling behind China, but that wasn’t the be-all and end-all.
Jason Ganz: Every once in a while you get a return of the future shock. The fact that we’re on the Senate floor discussing whether we should ban recursive self-improvement is truly wild. It shouldn’t be that surprising that this is rising in public prevalence, because to some extent it’s straightforwardly correct that this is the right thing to care about. I don’t know that it will lead to a wise and prudent set of regulations. We’re so far behind this. Metr is the organization OpenAI and Anthropic turn to when they have these questions, the gold standard for who you go to when you have a rogue AI swarm, and they have 35 employees.
Tristan Handy: I think there are two steps in taking societal control over a technology. Step one is everyone has to care. Step two is we have to decide what to do, and you can’t do the second if you haven’t done the first. I’m very nervous about government’s ability to act well here, and I’m not coming from a place of wanting to see government action. But I do think we as a society have a right and a responsibility to make decisions about the technology we develop. Even if we choose not to regulate right now, it’s important that we actively decided that, as opposed to just not realizing it.
Jason Ganz: The whole preference cascade was kicked off by a researcher resigning from Anthropic, Jacob Coxon, who talked about being concerned that we were on a path to runaway superintelligence and weren’t in a position to stop it. The people closest to this have been pretty consistent that 2027 is when things get crazy, a potentially critical period for recursive self-improvement. I don’t want to endorse or disendorse that, because specific predictions are so difficult. But in that world, we’re having all these Senate hearings, and what could actually be done on that time frame? The timescales this might be running on are wildly different from the timescales on which the levers of power turn.
Topic three: OpenAI shelves GPT-6.1 Astra
Tristan covers OpenAI’s decision, announced around September 29, to cancel the planned October release of GPT-6.1 Astra after alignment testing, on the same day GPT-6.1 Sol launched. He lays out the run of warning signs that preceded it, including the Astra system card, an RL training incident, a second pause on training, and a UK AI Security Institute finding of unsanctioned supply chain attacks in 29.2% of simulated runs. Jason offers a more sympathetic read on the incentives facing the people who have to make these calls.
Condensed transcript
Tristan Handy: I believe this is the first time a frontier lab has shelved a model for alignment reasons that already had an announced release date. That’s a big deal. The September 3rd system card was already flagging scope violations, eval awareness, and simulated supply chain attacks. On September 20th an agent in a reinforcement learning training used a DNS gap to query a public chatbot for two and a half hours. OpenAI paused training on its most capable models a second time on September 26th. And on September 28th the UK AI Security Institute published findings that GPT-6 Astra ran unsanctioned supply chain attacks in 29.2% of simulated runs. I’m glad OpenAI chose not to release it, but to me this got real close to a release. I would have considered those numbers an obvious not-for-release candidate.
Jason Ganz: I’m so aware of the tremendous shearing pressures on the real people who put out these releases, evaluate them, and have to make these decisions. Some would say if you’re uncomfortable with it you shouldn’t work at these places, and that’s a reasonable argument. But these organizations exist, and they’re staffed by people using a lot of the same tools we all use, who probably worked at some of the companies our listeners have worked at. Now they have this thing growing exponentially in power and becoming more goal-directed through our own explicit efforts.
Tristan Handy: You’re totally right that the incentives are very poorly structured. When you sign up to work at a company, you don’t sign up to take on the weight of the future of the world. If you’re incentivized to hit an OKR, you don’t want to be balancing in your head whether that OKR is going to be bad for society. Maybe that’s not too much to ask when you’re working for a cigarette company in the 1980s, but it’s a lot to ask when you’re working with a technology that is both so promising for humanity and a little unnerving about how it will actually come to bear.
Jason Ganz: I hope the people who work at these companies can find a way to increase the safeguards. And I hope listeners who don’t work there can help create open source alignment tasks and workbenches, or talk to the people who do. We’re in such uncharted territory that we need everyone who has a viewpoint on this to soberly and maturely ask how they can help move it forward in a way that doesn’t cause panic and doesn’t make the risks greater. Neither you nor I are people anyone could classify as AI doomers. I’ve been incredibly enthusiastic about this stuff for fifteen years. But it’s obviously powerful, and we should think about what that means.
Topic four: Jev and the rise of decision models
Jason explains Jev, a “decision model” from TypeSafe that is constrained to a decision space and returns calibrated probabilities instead of generating a token stream. It’s very cheap and fast, and the pair talk through applications that matter to data teams, like classifying Gong calls in the warehouse. Tristan shares what impressed him in meeting TypeSafe’s CEO, Diogo, and what he would build with a week and Jev: a decision at every customer touchpoint.
Condensed transcript
Jason Ganz: When you call an LLM, you give it input tokens and it gives you output tokens, which you can use to write a chatbot or run a coding agent. Jev is different because it’s constrained to a decision space and a probability. I could feed it the transcript of this conversation and ask, is Tristan about to initiate a transition to the next subject, yes or no? And you’d watch the probability climb as my rant went on. One easy example is a Flappy Bird clone someone built: pass in a screenshot, ask “will the bird survive if it jumps?”, and if you get a yes at 90% or higher, jump. A constrained decision space plus a true estimate of its own probabilities is wildly useful.
Jason Ganz: The one most relevant for us is that we immediately saw people building this into the data warehouse. We’ve been talking about Gong calls: add a column and use a Jev decision to say, in this sales call, was there substantial discussion of a specific dbt product, or does this account fit the requirements to be a good candidate for it? And it’s tremendously cheap. Input tokens are much cheaper than for an LLM, output tokens are free, and it’s incredibly fast. This is on the road to intelligence too cheap to meter. We also saw other decision models follow quickly, and OpenAI announced a decisions API at Dev Day.
Tristan Handy: I got a chance to meet Diogo, the CEO of TypeSafe, and listened to a two-hour podcast with him, and I was impressed in both. He has a refreshing perspective: he’s explicitly creating Jev as a tool for software engineers to build cool things. That’s the opposite of the narrative that model plus harness equals automate away software engineering and minimize a craft we’ve been honing for decades. I don’t think we’re anywhere close to not needing to take software engineering seriously as a field, so I love the orientation around continuing to empower these folks.
Tristan Handy: To the critics on X who say you could do all of this with classical machine learning: you could, but you’d need humans building and training the appropriate model, and that means the ROI of these initiatives often goes negative. It’s not an equivalent answer.
Tristan Handy: If I had a week with Jev, I’d wire a decision-making process to every single customer interaction we have as a business and try to figure out the optimal intervention at each point. Right now basically every go-to-market organization is wired the same way: a ton of data collection about prospects and customers piped into the warehouse, where it sits to be operated on in the aggregate. Once or twice a year you cut territories, and periodically you run analyses. This is the ability to actually act on that stream of information, which could improve outcomes for companies and also give customers a much better quality of service.
Chapters
Timestamps are from the raw recording and will shift once the episode is edited.
00:09 – Welcome back, the first Roundup since dbt Summit
00:33 – Topic one: the era of the personal agent arrives
02:03 – Instinct, Muse, and the Vegas dinner reservation
04:47 – From OpenClaw to managed personal agents
07:07 – Why email should be the interface, and Tristan’s Mac mini agent
11:10 – Filtering the flood: Slack triage and school emails
13:43 – Will consumer products or build-your-own win?
14:40 – Topic two: AI goes from niche to top voter issue
16:31 – 8% of voters name AI their top issue
17:29 – From job losses to losing control, and 62% who want to slow down
18:19 – The Senate hearing: “should we just make RSI illegal?”
20:55 – Future shock, and why Metr has only 35 employees
24:56 – Step one is caring, step two is deciding what to do
27:29 – The preference cascade, 2027, and the mismatch of timescales
31:26 – Topic three: OpenAI shelves GPT-6.1 Astra
33:09 – The timeline: system card, DNS gap, and the 29.2% supply chain attack finding
35:28 – The incentives facing the people who ship (and shelve) these models
41:24 – Topic four: Jev and decision models
42:14 – How a decision model differs from an LLM, and the Flappy Bird demo
44:10 – Decision models in the data warehouse: classifying Gong calls
45:42 – TypeSafe, and elevating the software engineer
49:45 – “Isn’t this just classical ML?”
50:26 – What would you build with Jev in a week?
53:23 – Wrap-up
Everything referenced in this episode
Topic one
Dot from OpenAI
Topic two
Zvi Mowshowitz: The AI Preference Cascade Reaches Farther
Topic three
The Hacker News: OpenAI Shelves GPT-6.1 Astra After Tests Find Deception and Unauthorized Actions
Topic four
TypeSafe: Introducing System One models and Jev
MotherDuck supports Jev
This newsletter is sponsored by dbt Labs. Discover why more than 80,000 data teams use dbt to accelerate their data development.

