Teresa Torres | How I Tested an AI Coach
Teresa Torres - Author, Speaker & Product Discovery Coach
In this episode, Teresa explains why assumptions look different depending on the level you’re working at and how using AI inside your courses can give students more opportunities to practice.
Have a Listen
Summary
In this episode, I’m joined by Teresa Torres. Teresa is an author, speaker and a product discovery coach who I’ve been a fan of for quite a long time.
We chat about why assumptions look different depending on the level you’re working at, why showing your work is so important for alignment, and how using AI inside your courses can give students more opportunities to practice and receive feedback.
We also dig into AI evals - how Teresa uses them to measure the quality of her AI coaching tools, identify failure modes, and systematically improve their performance instead of just trusting that an LLM will get it right on the first try.
If you want to learn more about testing assumptions, designing better learning loops, and using AI without giving up your critical thinking skills, this episode is for you.
Takeaways
AI can create a safer space for learning. Teresa found that people are often willing to ask an AI questions they might feel embarrassed asking another person.
An AI coach works best when it is grounded in a clear teaching model. Teresa’s interview coach was effective because it was built on years of refined curriculum, rubrics, and explicit feedback criteria.
AI can dramatically increase opportunities for deliberate practice. Instead of waiting for an instructor, students can practice repeatedly and receive immediate, personalized feedback.
Building AI tools can improve the underlying curriculum. When an agent struggles with ambiguous instructions, it exposes gaps in how the material itself is taught.
Evals are simply a way to measure AI quality. Teresa uses evals to identify specific failure modes, track how often they occur, and test whether changes actually improve the agent.
Domain expertise still matters. Recognizing that an AI coach has given subtly bad advice often requires deep knowledge of the subject, not just technical skill.
Product teams should define what “good” looks like for AI. Generic quality metrics are not enough; teams need acceptance criteria and evals tied to the unique value their product is supposed to create.
AI changes the role of the teacher, not just the tools. Teresa is moving away from being the expert who simply delivers answers and toward creating environments where people can practice, get feedback, and learn for themselves.
Guest Links
ProductTalk Website: https://www.producttalk.org/
Teresa’s LinkedIn: https://www.linkedin.com/in/teresatorres/
Transcript
David J Bland (00:02.658)
Welcome to the podcast, Teresa.
Teresa Torres (00:04.38)
Thanks David. I'm excited to be here.
David J Bland (00:06.518)
I'm so excited. I've been wanting to have you on for a while and I felt like, I had to actually make this a credible podcast first before I bring you on. Cause I I mean everyone looks at my stuff and they're like, yeah, but you're just testing it. You know, are you really gonna stick with it? And you know, we're fifty some episodes in. So like, okay, maybe Teresa will join. And I've just wanted you on for so long. Like you're just such a joy to hang out with, and I always learn from talking to you. And our listeners may not know, but you actually like
showed me around my wife and I around Portland. You were so gracious. we were looking and moving there and you're like, Hey, yeah, come up. I'll I'll help show you the the ropes. And we've just kind of stayed in touch online and in person and I just I I just really appreciate you making some time and maybe we can geek out over some assumptions and AI and stuff in our chat today.
Teresa Torres (00:54.798)
Yeah, I will say it's always a delight to chat with you. And I loved when you came to Portland. I was excited you might actually move there. Alas, it did not happen. but I moved away from Portland, so that's fair. and I would I, you know, there's a few people that I feel like there's just a super strong synergy for lack of a better word between the work that we do. so in a lot of ways I feel like you're a kindred kindred spirit.
David J Bland (01:03.352)
It did not.
David J Bland (01:22.008)
I appreciate it. Yeah. And we'll and we'll riff on some stuff about that. I would say I probably first found your work, I'm guessing it was probably like earning startup E days, where I had just kind of come out of startups and then I was pulled into consulting. And part of my consulting, I really had a hard time with people going, yeah, we're going to do a transformation. Here's your five year plan and just follow no, no, no, we're not doing that.
Teresa Torres (01:48.432)
Yeah.
David J Bland (01:49.742)
It's gonna be, we're gonna do some discovery. And then I got pulled. Actually, my first client, I think Marty Cagan had just came in and I gave a lot of advice to who we're talking to in a future episode. and then I was pulled in. I think it was myself and Jeff Patton and a few other people, Skip Angel, and and I was just trying to like figure it out. You know, it was like, my gosh, they're trying to organize around and discovery and assumptions and all this work. And it was so rewarding, but it was also.
Teresa Torres (02:09.074)
Yeah.
David J Bland (02:17.661)
Super challenging work for me because at the scale of the company I was in, and I remember coming across your work and I was like, okay, you have some tools and you have some ways of thinking that could help me out because I was just trying to, you know, tread water there and and just help them move forward, but not get bogged down and all this really, really heavy framework stuff. So I think that's when I first first stumbled across some of some of your your content.
Teresa Torres (02:45.948)
That's really interesting because I actually, you know, when I first started consulting and coaching, I went back and I got a master's in learning and organizational change. And it's because I thought this like organizational change stuff was so hard. And my takeaway from my master's was don't do anything in organizational change. And I really deliberately decided to focus at the team level.
Like I coach at the team level. All of my curriculum is at the team level. And it's very deliberate. It's not that the organizational change stuff doesn't matter. It absolutely does. But I'm so allergic to it. But I'm glad that like as you found yourself in that very challenging problem, it's just so hard. that you found my stuff helpful.
David J Bland (03:31.65)
Yeah, it was super helpful. And then I think just over the years, just seeing how your work has evolved, and all of our work's evolved, but just you you kind of are sticking to your principles, but you're very open to the idea of maybe there are new ways of delivering this value or helping people. And one of the things that I I don't see a lot of c consultants and authors and advisors do is they don't necessarily build in public. And I think you and I have talked about this in the past, but you're not afraid to do that. Well, why why is that? What is it
Teresa Torres (03:55.101)
Yeah.
David J Bland (04:00.255)
about you that you feel like, Yeah, I can try this in public and it might not be completely polished, but that's okay.
Teresa Torres (04:07.24)
this might be a little spicy. I I think there's different types of like consultants and advisors. I think there's some people that have found something that is inherently who they are. And I don't I'm not a consultant for the discovery habits. The discovery habits are intrinsically how I think. And so like I'm just being me. I'm just showing up and being me.
And that's not true for everybody. And I'm not saying that's like the one right way to do things. I think there's a lot of people that like sell the current trend. And that's a great way to have a business. but I also think it's scarier, right? Like it's harder to build in public because you don't have that depth of experience with it. And I think for me, like it's sort of a luxury of like, I just don't care that much about what other people think. Like I just this is who I am. This is what I believe, this is how I work. Like I
Feel like at my core what I teach is the scientific method. And the scientific method is really hard for humans to do because of all of our cognitive biases and our like feeling that we're so right when we're not. And I f to me it's like this game of how do we help humans get better at this and like teeny tiny nudges along the way. And the like how doesn't matter. The how is gonna change every day of our lives. But there's this like foundation that I just feel it like is how my brain works. Like I just
Wanna understand humans? I wanna I I'm very acutely aware of how often my perception of reality is just completely wrong. I'll I'll tell you a really funny story. I was this happened last night at dinner. I was sitting with my husband at like on a lawn in Bend, and he was wearing a tad Tab Benoit t-shirt. So Tad Benoit is a blues slide guitar musician. He's amazing. We just saw him live. And I'm looking at his t-shirt and I'm like, do you have two Tab Benoit t-shirts?
And he goes, no, I only have this one. I'm like, no, you have one that has like a vertical guitar with a cro al alligator on it. He's like, no, this is the only one. I'm so convinced even right now in my head that he's wrong. Cause like I can visualize this t-shirt. It doesn't exist. He literally has one t-shirt. And I can't like this is such a silly story, but it's such a perfect example of like how we can feel so certain that we're right.
Teresa Torres (06:30.896)
And we're totally wrong. and I love those moments. It's just such a h it's such a reminder of our humanity. And so I think because that's just like the thing I geek out on, the stuff I do for work is just me being me.
David J Bland (06:48.011)
Yeah, and you have a knack for explaining things in a way that people can get it. Like really complex kind of concepts and things. You you're you have a knack for kinda breaking it down. It's probably why we get along so well, you know, because we're we're we're alike in that regard. And
I mean me, I have a this may sound crazy. I have a hard time with uncertainty, you know, like just from childhood stuff and everything. And and people are like, What? You? And I was like, You but you teach people how to deal with that. I was like, Yeah, yeah, I do. I do. And there's a reason I do that, you know? and I got pulled up, like I feel like all these different trends I got pulled into, like agile early days and then the lean stuff and all this. I felt like, maybe here's a little process to help me deal with uncertainty, you know? But what I noticed over the years, it feels like
Teresa Torres (07:16.018)
Yeah.
David J Bland (07:32.32)
It also draws a lot of other people who have a hard time with uncertainty. And then we project so much onto this way of working that we just overcomplicate it and try to eliminate all the uncertainty. And I've been pretty good. I think lately I've been skewing back towards, hey, principles, I'm not so much about frameworks anymore. I just want you to, you know, think through things and look at the risk and try to address it. And
I think it's just in our human nature, maybe just to overcomplicate and keep adding more and more and more over the years. And and that's something I'm really trying not to do.
Teresa Torres (08:07.568)
Yeah, you know, so first of all, on your comment about not being comfortable with uncertainty, I'll share for me, like mine is I really like being right. You know, like I talk a lot about leave room for doubt. I really I'm I'm one of the most confident people you'll meet. Like I I am very certain I'm right. And I think this is partly why I embrace the leaving room for doubt. Like I know this is something that I have to do. Like, there's
I can't remember what it is. Maybe it's the Hogan assessment. There's a h assessment in the like leadership and development space where it talks about your strengths are your weaknesses. They are one and the same. And so there's an attribute about you that like is a is a huge strength in some contexts, but it's also a huge weakness in other contexts. And for me, I feel like it's this, it's this yin and the yin. How do you hold both? I really want to be right, and that's what gives me a bias towards action and allows me to act. And I have to leave space for doubt so that I test my ideas and test my assumptions.
And I think that's the fun of it, is how you hold those competing ideas in your head. So, like, be really uncomfortable with uncertainty and also get comfortable with uncertainty.
David J Bland (09:16.875)
That's a great way to frame it. at least into a topic I do want to explore and I think all our listeners want to explore is around assumptions. And and I'll just start with kind of my background on it, and I'd love to know more about where you landed. When I was doing this early work, I was trying to figure out how to design experiments, and and my teams would just get really excited about experiments.
But then they would test things where we're like, hmm, we're not really reducing risk at all with this. I mean, it's kind of fun. Maybe they're learning a new type of experiment or a new tech stack or something. But I kept coming back to, okay, we have to kind of pull this back to our strategy and our risk and and try to de-risk what we're working on. And I tried a bunch of different things. I tried stack ranking, I tried dot voting, the dotracy, right? I didn't think it should be a popularity contest for which assumption to test.
And finally land landed on this two by two, which I've changed at least three times since I've been using it. And it was kind of a mix of, well, I was exposed to like Jeff and Josh's work back at Neo. And then I was exposed to a lot of design thinking, which you have a really strong background on. And I just started blending it together. I feel like sometimes I'm more of an editor than necessarily like a tool creator. And when I was like, okay.
I think this two by two can help. Wait, my my teams are already thinking and kind of these three themes. Let's blend that together, which eventually became, you know, my version of assumptions mapping. There's many different versions. And I use it as a way to get a team just to talk to each other about risk and prioritize it. But I feel like people fall into these camps of like there's a right and a wrong way. And like how are you approaching it now? Or what's your journey and how you're trying to get people to prioritize risk around, you know, assumptions?
Teresa Torres (10:58.652)
Yeah, so I also feel like an editor on this because I borrowed a lot from you. I borrowed a lot from Jeff Patton. I sort of tweaked it to make my make it my own. I feel like we you and I work in at different scopes. and so I sort of modified it to work better for the scope that I was working at, which I'll explain that. I think for me, what attracted to me this to this space is I always have had these experiences like working with a team where
In my head, it was really clear the scope at which we were working. Like we're testing this thing right now. But there was always these like disjointed jarring moments where somebody would throw out a suggestion that had nothing to do with the thing we were trying to do. And in my head, I was like, you're insane. And what I didn't realize was that like we just weren't working from the shared scope. And this is exactly why I like started to develop the opportunity solution tree. Like, we're working on this outcome, we're working on this opportunity.
We're working on this solution. We're testing this assumption. It's all about like how do we stay aligned as a team. And so I think this idea of like where is there doubt in an idea? Like where it you talk about it as risk. I framed it as doubt because I saw this amazing quote by an organizational psychologist, Carl Weck. He said, wisdom is having the confidence to move forward, the confidence to act.
Well maintaining room for doubt to know you're you might be wrong. And I was like, wow, that I get goosebumps still from that quote. Like, what a great combination. And so I was great at the have confidence to act. And I knew I had to buffer the like, okay, but leave some uncertainty, leave room for doubt. And the what's hard about that, it's hard enough to do that as an individual. It's really hard to do it aligned as a team. And so
I really nerd out on like team collaboration. How do we set the scope so everybody's on the same page? And I wasn't great at this at the solution level until one day I was at a collaborative gain event and Jeff Patton was doing a workshop. And I saw him story map in person for the first time. And I
Teresa Torres (13:15.494)
You know, to be honest, I'd never worked at very big companies. I'd never really seen like Agile and its full glory in practice. I'd written user stories, I ran sprints, like I knew the ceremonies, but like I had never heard of story mapping. And I always I think I'd heard of it, but I heard of it as this like a release planning thing. But when I saw Jeff do it, I was like, you are user journey mapping, a user's experience through a solution. And like a light bulb went off in my head. And what went off in my head was
This is doing it's aligning people around what the solution, what is required of the customer to get value from the solution. And then two, it's forcing you to get specific enough about your solution that you're making assumptions. And like within five minutes of watching Jeff's story map, I was like, this is how I'm gonna teach teams how to generate assumptions. So this is where I think our scope is a little different. You like to generate assumptions like at the business model canvas level, right? Which I love.
And then I get teams to generate assumptions at their specific solution level, and I have them tie the assumptions to a single step in the story map. Because that makes them easy to test. They're now scoped appropriately for a fast test. And then we the part I borrowed from you is I love your mapping of desirability, viability, feasibility to the canvas. And I was like, I can map that to story map steps. So
For every step in a story map, you have to want to do that thing. You have to be willing to do that thing. You have to understand it. You have to be able to able to do it. These are our desirability and usability assumptions. And I can just have the team start to enumerate. these are all that here's where there's risk in your idea. And then for assumption mapping, I got a lot of flack from people who had learned it from you, because I teach it very differently. I tell teams you cannot talk while you do this exercise.
And people are like, no, David says the whole point is the conversations. I'm like, yeah, when you're doing this at the business model canvas level, like as a founding team, you should be aligning on these things. I go, but as a product team, you're working with three ideas right now. You're probably gonna build none of them. They're all gonna get thrown away, because they're all dependent on faulty ideas, assumptions. And what I want you to do is do assumption map as quickly as possible, so you start testing, so you start throwing away the ideas. And
Teresa Torres (15:42.053)
This came out of me teaching in-person workshops and seeing teams just get stuck on debate. And at the solution level, it doesn't matter what assumptions you test. It just matters that you start, right? So like I don't have to optimize for the riskiest in a business, I need the riskiest idea in the business. In an idea, I just need to start, right? And I think teams don't realize they're gonna throw away half a dozen, a dozen ideas before they even circle on something that might work.
And so I just want them to like have super fast cycles.
David J Bland (16:15.946)
Yeah, I appreciate that. And I think at different levels, even in my work, right? you know, I I did a lot of stuff with Alex Osterwalder. That's why I wrote a book with him. And when you think of that tool, it kind of plugs into canvases really well. And that was s purely out of a need in my coaching.
Where I was like, these people got all these canvases. It's kind of like innovation wallpaper, they're not doing anything with them. I need to have them remind that we did all this work here for a reason and it ties into your testing. And I just kind of was able to plug into that. And he and Eve had been kind of layering the DVF on the canvas too. So I can't take all the credit for that. But I think our book was the first time we published a book with that framing. And even that, it's not perfect. I mean, if I had the nitpick, I'd say.
Teresa Torres (16:36.934)
Ha ha.
David J Bland (16:58.068)
Is a channel really desirability? You know, like like it it it kind of looks cool from a, you know, a really simplistic level. But when you dig into it, you could probably make the case that maybe the boxes could be colored a little differently and everything. But my point was don't treat that thing as a series of facts and also just don't ignore it after you've built it and then go test and then have to recreate everything for your test. And so I think with our scope.
Teresa Torres (17:03.261)
Yeah.
David J Bland (17:23.978)
Even if you look at canvases, the business model, you're gonna have assumptions at the overall business model. But if you did like a value prop canvas, you're down in the weeds on the jobs, pains, and gains of the customer. You know, you're down at that segment level. It's a different level of assumptions. And I think I could probably do a better job of communicating that to people of depending on where you do your mapping or what you use to extract from, you're going to have different levels of assumptions. And I've been trying to open it up too. So it's like you can extract risk from canvases, but you could extract it from your PRD.
could extract it from your roadmap you could extract it from your exec one pager and so when I'm talking about extract map test I'm trying to open it up a bit because I think people kind of read the book and they go I have to have a canvas is like well you could but you don't have to you could use other things as well and the levels are going to depend on what you're extracting from
Teresa Torres (18:10.822)
And yeah, and I think there's different stability at different scopes, right? So hopefully your business model isn't changing all the time. And that's where I would say you're right. It's like slow down, have a big conversation about this. Let's get the risk right. Let's let's test more of these assumptions. Then you might move down to a value proposition. That's might change more often than your business model, but isn't changing as often as like feature ideas.
And then there's a lot in between, like you said, your roadmap, like that's changing more than your value proposition, but not as often as your feature ideas. And so I think there's this, I don't know, this is how my brain like just works. I like to think about it as like different layers, different scopes. How do we match the activity to that layer or scope? and the time investment. Like I don't wanna invest a ton of time in an idea that I'm just gonna throw away.
David J Bland (19:02.828)
Agreed. And I just want to pause a moment and say how amazing Jeff Patton is. Like I learned how to story map from him as well. And I think you and I probably came to the same realization but different ways because
Teresa Torres (19:06.93)
For sure.
David J Bland (19:14.292)
I would look at that map and I go, there's assumption here and there's assumption there. And there and and we literally just laid them out. and I remember him and I arguing, well, maybe just debating on I think I took it too far because I looked at a team's backlog and I was like, You should have an assumption on every story. And he's like, No, don't do that to them. And so he and I went back and I'm like, Well, wait, maybe you need to rain me back a little bit on this. And he's like, Look, they're trying to log in.
Teresa Torres (19:17.309)
Yeah.
Teresa Torres (19:31.278)
Yeah.
David J Bland (19:38.656)
They already know how logins work. You don't need a big assumption on login. Like, and so he did like kind of level set me a little bit, and I appreciated that from him. But I have to say, I'm this is a Jeff story. I watched him facilitate once and I was just like in awe because he had a group of people and he was trying to get them to say whether or not they agreed with statements he was making. And he just did like a fist of five, and we're talking like a group of fifty people in his company, and they would
Teresa Torres (19:52.85)
Yeah.
David J Bland (20:06.527)
Hold up their hands, like you know, one to five, right? And he would turn around, he would look, and you could see the gears turn a little bit. He would turn around and draw a pie chart and he would fill in the pie chart on the statement based on the number of fingers he saw. And he went through a series of statements and then he looked at and he goes, Okay, here are the ones you agree with more, and here are the ones there's like some division. And I was like, my gosh, like I c how do I get to that level of facilitation? So he I've learned so much from him. He's I have to get him on here at some point because he like
Teresa Torres (20:30.322)
That's amazing.
David J Bland (20:35.881)
He's just he's such an interesting person to be around and he has a style where you're like, my god, I've never thought about doing it that way. And it's I've just learned so much from just being around him.
Teresa Torres (20:45.436)
His style is so uniquely him. I think that's what I love about it. And he he doesn't it's easy to underestimate him. He does not look like he's gonna hold the whole room. But I have watched him hold the whole room every second, right? It's so cool. I love his work.
David J Bland (21:03.721)
Yeah, I have to say I'm not brave enough to do document camera and draw from scratch. I defaulted back to, I can annotate over slides I have in in place. And that's the limit. And I have a design degree, and I'm still not at the I'm going to draw from scratch point of view. But he it sounds like he was influential influential to both of our kind of thinking. And when I when I think about your work, something that really stood out to me, and I think this was at Mind the Product in San Francisco, where you're making this point of just like, hey, show your work.
Teresa Torres (21:11.004)
Yeah.
Teresa Torres (21:17.788)
Yeah.
David J Bland (21:33.341)
Show your work. And can you explain maybe a little bit of like the reasoning behind that and why it's so important? And then maybe we can talk about how AI might be changing that. But like this concept of show your work, maybe you can explain a little bit to our listeners what you what you mean by that in your context.
Teresa Torres (21:50.301)
I think this is actually really rooted in my own personal experience in that it's related to some of the things we've already talked about. I think I naturally like go deep. Ma I will nerd out on something really deep. And so, like as a product manager, I would go really deep on I'm gonna design this perfect solution, I'm gonna learn all about my customer. And then I would go and do that by myself. And then I'd have to go convince executives, and I would just say, here's what we should do.
Well the challenge with that is like I didn't show any of my work. They have no idea how deep I went. They had no idea I even talked to customers. And when that's when you don't show your work, like they would just react to that and be like, okay, that's one idea. What about this other idea? And I'd be like, that idea is so shallow. Like, what do you mean what about this other idea? Like who did you talk to? Like what did you do? And I didn't really realize that like the I wasn't being persuasive because I wasn't showing my work. I was just presenting my conclusions.
And then as I started coaching teams, I saw that like we all make this mistake. And so we talk a lot about like in like we talk about product managers influence without authority. Like authority allows you to influence. Let's be real. Authority does not really allow you to influence. you still have to learn how to influence. And so I just started to realize like some of this is grounded in a fallacy. Like there's this fallacy
That we have that all humans have of like, if only you knew what I knew, you would agree with me. And so there's a little bit of that in the show your work. But but if we don't if I don't show my work, it's very unlikely we're gonna agree. If I do show my work, it increases the chances we're gonna agree. And more importantly, when we don't agree, you can see where in my work there's a problem. And we can have a much better conversation. So you can say,
Like that's great you talk to these three customers. I talk to these three customers and they have a different need. And now instead of like debating about our conclusions, we're actually having a conversation about the foundation on which we're making conclusions, which is way more powerful.
David J Bland (24:01.653)
Yeah, I like that. I like that. And and I'm wondering with showing your work, I've been watching you kind of from the outside 'cause we haven't really caught up recently like this, but your work with AI and evals and all this all this interesting stuff that's going on in our industry right now. How does show your work apply there? Or what do you see changing at all, if if it if anything?
Teresa Torres (24:25.244)
This again is also pretty personal. I okay, so for 20 years I've been working as a discovery coach, 15 years full-time as a discovery coach, but lots of years developing teams before that. And I love the work. Like I genuinely love the work. I love writing, I love teaching. This is intrinsically who I am. I hate being seen as a guru. Like I hate it. I hate going to conferences and having people like be embarrassed to meet me.
I don't mind taking pictures, like that it that's fine. But I hate the like you're not seeing me as a human, you're seeing me as a caricature of a famous person. I hate it. and I've really come around to just some humility around like what is my job as a teacher, and it's not to impart knowledge on you. My job as a teacher is to create environments that support your learning.
And to me, that's a very important distinction. One is this like expert model of like I will impart my wisdom on you. And the other is no like, no, like my job is to create space so that you can learn. And I think this is coming out of lots of years of teaching and seeing a lot of variation in who learns and how well they learn. And so when AI happened and I started learning about it and digging into it, first of all, I just started having a ton of fun.
David J Bland (25:25.195)
It is.
Teresa Torres (25:50.941)
And then people started asking me, like, why won't you why like why aren't you creating AI courses? And I was like, because AI courses would just be more of me being the expert imparting my knowledge on you. And I want to retire from that. Like I don't want to play that role anymore. And I don't mean that like I don't want to teach and I don't want to help. I just don't want to do it that way. And so I've been really playing with how do and then also there's all this noise. Like there's so many snake oil gurus out there.
And like I don't want to be in that category. Like it's just disgusting to me. And so like I started to really think about how can I show up differently and how can I use this new it's almost like I get to start over. So like what would I do differently this time? And so like I'm very deliberate about I'm not gonna tell you how to do things. I'm gonna show you how I'm learning. And then I'm gonna create space for you to show up and connect with each other and do the same thing.
So like instead of creating an AI course, I host AI Maker Studios and I host AI show and tell. And it's all designed to get people to come together and learn from each other. And like I c I share too. I don't think I have nothing to share, but I don't I'm trying to break this model of like Teresa the Expert. It doesn't it doesn't feel like a comfortable hat for me anymore.
David J Bland (27:07.007)
Yeah, I I can I can relate to that a little bit. I think our AI journey, I mean, I I basically I know you have a history even before teaching, if I remember correctly. Like, you know, back to your traditional schooling even. From for me, it's more of a like I started as a designer slash software engineer at startups. And then quickly it's funny, my first startup I joined, I was doing design work and they're like, Can you just cut all that up and make that work on the front end? And I was like
Sure. Like please don't make me go back and like live with my parents. So I just learned it, you know. so I I can't say I learned the greatest habits as a software engineer. But and then got into API development for a while. And then my career just kind of took a turn. I think it was like the third startup I was at was imploding, and I was carrying my stuff out in the box and I was talking to my wife on the phone. I'm like, hey, honey, I'm coming home. And I was like, I cannot keep doing this. And I started doing like advising and consulting, and that was around 2011 or so. And I was like, wait.
I have a knack for explaining things to people. Maybe I run with that and then I kinda yeah, I could dabble in you know, coding here and there. But I have to say, like when the AI started coming back with, you know, this like mostly GPT stuff and I was learning prompting and then custom GPTs and then agents and then kind of playing around with evals, although I want to get more into that with you. It it's just been a really interesting journey of like, okay, I can I can talk about this publicly and kind of reason through what I'm working on.
And I'm not afraid to do that. I don't have to be so polished that I don't share like how I'm learning. And and my learning has changed quite a bit over the last three years. But the one thing I've not really sort of fully got my head around yet is what I'm observing is almost like a interesting dynamic between like agents helping people with learning and how they receive feedback from that versus people.
So I I I mean I'll share this with you. So so basically what I'm seeing right now is like I can give advice. I can say, hey, your map should be like this and this, or have you thought of this? Or the stakeholder can come in and say, hey, have you considered this? But if I put that rubric in an agent and the agent's giving that feedback, the way they take that feedback is different than a person giving it because it doesn't have all that baggage. Are you seeing stuff like that? Or are you playing with stuff like that? Or like what's your experience right now with just kind of blending the stuff into your work?
Teresa Torres (29:20.604)
Mm-hmm.
Teresa Torres (29:28.506)
Yeah, I love using AI in my courses and you and playing with how does AI change how we teach. before I get into that, I want to highlight we have really parallel paths. I started it as a designer that also did front-end development. I only worked at early stage startups. I started consulting in 2011. so almost exactly identical paths, which is sort of funny. actually, I think the first time I played with AI beyond just using it as Google.
Like the first time I had a real conversation with ChatGPT, it was actually a very politically charged conversation. It was on October 7th. And Hamasid just attacked Israel. I was really embarrassed by how little I knew about the Middle East. And I sat down and I just had a conversation with ChatGPT about it for about four hours. And I asked like every embarrassing question that I would never ask a human because I'd be afraid of putting my foot in my mouth and like offending somebody. And
I didn't have a lot of trust. Like I was like, I need a source for that. Like, how can I go learn more about that? Like, is that really true? And I asked like anything and everything from like even at some point in the conversation I was like, wait, I don't understand this term anti-Semitism, because like semitism has a meaning. And then anti-Semitism is used in a way that isn't tied to the meaning of Semitism. And I was like, explain, right? And so like we even got into like origins of words.
And then I was like, okay, so why is Jordan different? Like Jordan is next door. Like what makes it, you know? Like I just I would have never admitted to a human. This is now years ago, so I ha have no embarrassment over it anymore. Like I never would have admitted to a human that I didn't know these things. And so it was really empowering to ask all my embarrassing questions. And like I think I'm a good enough critical thinker. Like I had some awareness of like, okay, this is a political topic and
Chat GPT might be biased, so how do I like argue the other side? Give me some sources, I'm gonna go look into this. But it was like I had a very safe place to explore a very politically charged topic. That was very eye-opening to me. And as an educator, I was like, wow, if we could get people to do this, this could be a game changer in how people learn. And so I started to play with like ha like.
Teresa Torres (31:52.871)
This was months and months. It may have even been more than a year before I built my first teaching AI tool. But it planted the seed. And I started to like really this is what convinced me to start using AI every day. Like it really just forced me. I'm like, okay, I'm gonna explore all the boundaries of this new magical box that we have. And I just started using it for everything. I started using it for all sorts of work. And then eventually that led to
Okay, the way that I teach, I have a very strong belief in deliberate practice. I have a lot of long strong belief in we learn by doing. I'm a big believer in the I do, we do, you do model. So first I model it, then we do it together, then you do it on your own. And I was like, and in all of our courses, we d had live cohort-based courses where students came to class, the instructor modeled it, they practiced it, and then they and they gave gave each other feedback. And
students aren't very good at giving each other feedback, and they're not very good at being beginners at something. So they're like embarrassed to be a beginner, they're embarrassed to give each other feedback. And we would do a lot to overcome this. We would talk about feedback as a gift, you're not gonna be great at it to start, it's okay, like we'll we'll get through it. The instructor would give better feedback. And so this was the area where I was like, okay, I bet AI could give really good feedback. So then that was my first AI product, was I started to look at can AI grade your interview transcript?
And this is I built this in three weeks. And I that same exact version has been running in our produ in our course for two, one and a half years. and it's really good. I still read every single coach response because I'm afraid that it's somewhere it's gonna give bad advice, and it's really good. And what that means is my students now can have infinite practice. They can practice as much as they want.
David J Bland (33:23.262)
Wow.
Teresa Torres (33:47.891)
And they can get very personalized detailed feedback on their practice. That's pretty cool. There's another thing that's really cool. Before I had no visibility into how much better my students were getting. I had to just pop around the breakout rooms and see like week to week are they getting better. Well, now I have a score for every interview they do. And I I have four scores, so I can see where they're getting better and on which dimension, which is really great feedback for me on where does the course need to improve? Like what are people not getting?
So this was like my first teeny tiny step and now I have AI teaching tools in every s in almost all of my courses.
David J Bland (34:24.861)
That's amazing. Yeah, I remember I think you were just building your coach and we talked a little bit. I didn't realize you you did it in in in three weeks, which is pretty impressive. But you also have a lot of opinions about how it should work and the kind of rubric you should have inside there, which I think is hard to replicate, right? That's a lot of from your experience learning that over the years. I've done somewhat similar things with
kind of assumptions and assumptions mapping and my experiment library is now in like a JSON CSV format where it has all the sequencing and everything where it could plug in. And I've been noticing little things like that change the way I facilitate even. So for example, I would say at first I'd say, you have your three circles, you have to write down your assumptions for each circle. But more and more lately I'm just like just write you know I'll give them a framework sure for writing, but then we can feed that into an agent and it can help you categorize them.
we can also throw in supporting docs and it can extract stuff and say where your blind spots are. And it's not the same as me giving, I mean, I can give that advice, yes, but it's it's interesting how they take it from something that they think is neutral. It's not completely neutral because it has all my biases baked into it. But it's just so fascinating from a human interaction point of view to see how people respond to it. And and I think I'm sure you're s you're learning over the years too, as you've been doing this a little longer than me, like how it's
shaping how people how people learn.
Teresa Torres (35:52.167)
Yeah, so the first thing I'll say, the only reason why I could build this product in three weeks is because I've been teaching this class for nine years. Right? So for nine years with real humans, I've refined a rubric. And I'm all you for people that have read my writing, this won't surprise you. Like, I'm a very structured instructor. Like my classes have a rubric. I'm teaching against the rubric. I'm evaluating students against the rubric. We share the rubric with students. For people that aren't familiar with rubrics, it's literally like I teach.
a structured form of interviewing. An interview has four components. We're teaching exactly how to do those four components. And so when it was time to build the interview coach, all I did was I took the same material I used to teach my students and I gave it to the agent. Now it didn't just work out of the box. I had to create evals. There were the agent did make mistakes. I had to figure out how to measure those mistakes. I then had to figure out how to run some experiments to reduce those mistakes. It wasn't just like
I just wrote a prompt and it worked magically. There was more to it than that. and that didn't all happen in three weeks. The evals came later. but after three weeks, it was like in my class and being useful. and so I think one thing that I have realized is like educators are probably gonna be our best AI collaborators. Like, the way we teach humans is the way we teach agents. And I'll share.
So I have an interview coach, I now do AI generated interview snapshots, so customer interview synthesis, I do AI generated opportunity solution trees. And the more I build with AI, the better my courses get. Because the AI shows me where there's ambiguity in the way that I teach. So my teaching has to get really sharp for the AI to be good at it, which means my courses get better. And so what I love is I now like, I created like a domain knowledge library.
Where I'm just like it I build prompts for agents from that library. I build course material from that library. And like I almost can use my agent work to test how good is my curriculum. Because if the agent can't do it, a human probably can't do it. And so it's like my job now is just to make this domain knowledge as good as possible. And then the way I sell it is through my agents work and through my courses. And it's really fun.
David J Bland (38:13.373)
That's amazing. I love how you're approaching that. I think yeah, similarly, it's interesting how people look at like the book I wrote and they're they're like, You kind of sound like AI. I was like, look, I spent a year applying a taxonomy to forty-four experiments and there's two hundred pages of a very specific structure on cost, evidence, strike, setup time, runtime capabilities, DVF themes, and how to run and like like it's just
Teresa Torres (38:28.423)
Yeah.
David J Bland (38:39.485)
It's so interesting how many people I I I've I've spoken to over the years that said, yeah, I wanted to write that book, but I I didn't. I'm like, Yeah, there's a reason you didn't. It's super hard and you would have to focus for a super long time. So I do have a knack for like digging into things that kind of other people don't want to do and figuring it out, which is maybe a separate conversation we can have someday about copilot, which is I'm digging into because all my clients are locked into copilot. but something I want to pick out of your answer there is evals.
Teresa Torres (38:46.46)
Yeah.
David J Bland (39:05.357)
And I don't wanna go I know we can't do super deep and we can't teach evals on on you know a ten minute conversation, but can you explain a little bit like what you consider an eval and how it's helping you with with your work?
Teresa Torres (39:18.736)
Yeah, I think the best analogy for evals is just behavioral analytics. So when you use a product, a normal regular deterministic code product, we're tracking all sorts of things and we can see what you did and when. When we have an AI product, you know, I see people on LinkedIn write things like, how is your customer supposed to trust a non-deterministic AI that you built? And there's this like assumption in that comment that the AI is completely unpredictable and we can't guarantee quality.
That assumption is wrong. We can guarantee quality, and the way that we guarantee quality is with evals. So we steer the way the agent works in a number of ways. We steer it with how we prompt it, we steer it with what context we give it, we steer it in how we orchestrate the steps that it takes. So most AI products are not just one agent call saying do this task. they are a sequence of calls.
Where you can run deterministic code to check their output before you call the next call. So AI products are more complex than like you having a normal conversation in a chat chat GPT window. And the output, no matter what the complexity is, you have an input and an output, right? And we can grade the output. So an eval is simply a way to grade the ep the output.
The way the reason why I equate it to behavioral analytics is that eval is just a metric. That's all it is. It's just a metric. How are we going to measure a thing? And so the question is, is what do we measure? And the way we figure out what to measure is we have to look at a lot of outputs. We have to look at a lot of input-output pairs, and we have to say, did the agent do the right thing? And if the agent didn't do the right thing, we have to get really specific about what did the agent get wrong. And I'll give some real examples of this.
And then once we identify what the agent gets wrong, we can define an eval to measure how often does that happen. And then we have a just like we have a suite of behavioral analytics, we have a suite of evals that's measuring how often does this thing happen. And then once we can measure something, we can start to work on: okay, do we change our prompts? Do we change our orchestration? Do we change our context to reduce that error rate? So when I built my interview coach, it would make some mistakes.
Teresa Torres (41:37.617)
Inevitably as any LLM would do. And one of the things it would do is so I teach story-based interviewing. I want you to keep the participant grounded in a specific story about their past behavior. And the coach would like s give you feedback on a question you asked. And it would be like, You rambled a little bit, you could have been more concise. Try asking, tell me about your typical day. Okay, no, that's bad.
That's it, that's not a specific story about your past behavior. That's a general question. And that's the exact opposite of what we teach. That's a error. That's a that's an error. The agent made a mistake. So I can identify that. I can say, how do I build an eval? How do I create an eval that evaluates how often that error occurs? And actually, my eval for that was really simple. I looked for the words typical, typically, usual, usually, always, never. right, like there was like a handful of words.
That was a pretty good indicator that you were suggesting a question that generalized. And so now that's just code. That's just deterministic code. So when my interview coach runs, I can run deterministic code that can tell me yes or no, did this coach suggest a general question instead of a specific question? That's an eval. Right? So as you build AI products and you start to see the errors the AI makes, you can build up a suite of evals.
David J Bland (42:51.958)
Ooh, I like that.
Teresa Torres (43:00.188)
You can run them on all your input-output pairs and start to get a sense for what's your error rate. Now, once you know your error rate, you're set up to run experiments. Okay, I'm gonna make a prompt change and try to get the AI not to suggest a general question. And once I make that change, I can run a bunch of input-output pairs and I can run my evals on them, and I can see did I reduce my error rate? Right? So now I'm now I have like a perfect experiment harness.
now that I just described that in a very simple way, this gets very hard. The hard part is you identified an error, how do you measure it? So that's the first thing that's hard is how do I get a good measure of this error? the second thing that's hard is like, okay, I can measure the error. How do I reduce the error rate? and that's sort of the like product design part of it. What fascinates me about this is this does require a lot of domain expertise.
Like I wouldn't have caught that error if I wasn't an expert in interviewing feedback. It also requires like measurement expertise. Like, how do I like no engineer who doesn't have my domain knowledge is gonna figure out we can just look for these red flag words. Right? But it also requires engineering. Like you have to write code to measure that. and so this is a really interesting space because I think it's gonna require.
domain knowledge from product managers and designers. It's gonna require engineering skills from our engineers. I think coding agents can scaffold it a little bit so product managers and designers can do more of it. But it's this is like a new skill we're all gonna have to learn. And it's very cross-functional, which makes me love it.
David J Bland (44:44.487)
Yeah, it comes back to cross-functional. I I love that explanation. I feel as if maybe this is somewhat weird to say, but like my first introduction to just iterative and agile and everything was like through XP, like extreme programming. And over the years it's like, here's how you do T D D and then and then here's how you do B D D. And of course, most people did not do these things. And I feel like are we in are we in a place where AI is gonna like actually
Bring T D and B D D back? Because what you're explaining is almost like what's the acceptance criteria for this thing? And if you don't give it acceptance criteria, you are not going to get the quality you want. I mean, in a weird way, are we sort of gonna see AI help us be more disciplined in that regard?
Teresa Torres (45:31.398)
I think we already are. So a lot of people that use coding agents do start with tests. So I think like for the first time in human history, test event driven development is actually possible, right? because the agent because like nobody likes to write tests. Everybody likes to write code and to solve the problem. And so like it's like this ridiculous amount of rigor and discipline to actually write the tests first. but Claude will write the test for you first. No problem.
David J Bland (45:58.909)
I remember doing red green refactor at this one startup I was at, and I just watched all this crap break and I'm like, I didn't want to see any of this. Like, I mean it's good. But I was like, man, I see why people don't want to do this. but the whole like the idea of ex like evals and an acceptance criteria, I think maybe we need to make that linkage more explicit or just make it really clear to people that you know, this this is why it's going to help.
Teresa Torres (46:06.512)
Yeah. Yeah.
Yeah.
David J Bland (46:27.068)
And I just think it's really interesting that you found words that were kind of signs. And and I don't know if that's going to be very common across, you know, all of our, you know, peers when they start doing this stuff. But I love that you were able to recognize here are the words I need to look out for.
Teresa Torres (46:46.066)
There's okay, so the interview coach has I think seven evals. So I've identified seven pro prominent error classes that I had to measure and then reduce. which is not a lot, to be honest. Like it did not take very long. I was very surprised at how quickly I could get the interview ki coach to be stable and of high quality.
But I also do AI interview synthesis and AI opportunity solution trees. That is much more complex. In fact, I'll share I spent three weeks just recently. In fact, next week I'll be releasing a blog post about this. I spent three weeks chasing one just trying to get an accurate measurement of one failure mode. Just measuring it. Like not even fixing it, just measuring it. and it's it's tricky. And so there's
The I think the reason why there's a lot of confusion about evals, like I want people's takeaway to be an eval is simply a metric. How often did an error happen? That's all it is. The reason why there's a lot of complexity behind it is there's a lot of ways to do evals. So we can just write code. A code assertion is did this thing happen or not? So my looking for words is a code assertion. It's just regular deterministic code. Does the output contain this word? Yes or no?
but a really common way to do evals, like the most common way to do evals, is people will create a data set. They'll create an input, an output, and a s human labeled score. And then they'll run that input and output in the LLM, and they'll compare the output to the human labeled output, and then grade the LLM. And that's like where almost everybody gets started, and people talk about that date that golden data set eval as like a new product management skill.
Your job as a product manager is to generate input-output pairs that represent like it's acceptance criteria. Like how should the agent behave across these realms? And then they give that set to their engineers, and the engineers like iterate and try to get it to match the right score. The challenge with that, like that's a great place to start, and it's where almost everybody starts. But it's rife with assumptions. Like, how do you know what the right input-output pairs are? How do you know what your customers are actually gonna do?
Teresa Torres (48:59.334)
It's something you to maintain forever. Like forever, right? Like, is it always growing as use cases grow? Are you making sure it represents what you see in production? So there's a little bit of this like burden of maintaining this data set. And I swear, if our job as product managers is just to do these input-output pairs, like sorry, I'm done being a product manager. Like that sounds really boring. I think these code-based evals and then LLM is judge evals, where another you design a task.
where another agent evaluates the output of the first agent is much better. These scale much better. You design these based on errors you actually see the agent make, not errors you assume the agent will make. and I think both of these, like coming up with a strategy for how code can measure it, not writing the code, we can let a coding agent or an engineer do that. But like coming up with the strategy for how code could do it and then coming up with the prompt for the LLM as judge.
I think it's very much in our product manager's ability.
David J Bland (50:00.295)
Wow, that's much more nuanced answer than I was expecting. I mean not because of you, but but more of just like the nuance of evals. And something that you said also terrifies me about product is I already see product being I manage Jira tickets. And it's not such a big leap to say, I manage inputs, outputs. So I'm really hoping people are listening to this.
Teresa Torres (50:20.508)
Yeah, yuck.
Teresa Torres (50:25.244)
I'll share from a quality standpoint. I think the input-output data set eval is really limited. Like it just assumes you can predict what your customer is gonna do. We can't predict what our customer is gonna do. And I know some companies like they actually sample production traces to update these data sets, and that's a little better. But I really like.
Come on, we're product people. We can be smarter about this. Like we can think about it from an acceptance criteria and like what do we think the quality bar is and how do we define evals to measure that that metric of quality?
David J Bland (51:00.806)
I I agree. I think I keep coming back to acceptance criteria because that seems like a familiar thing that was somehow maybe neglected in the push over the years of how things have changed and how we strayed from kind of the structure around things. And it's it would be really interesting to see what product looks like over the next, you know, five, ten years. And I I'm really hoping it's going more in the direction that you're describing and less in the direction of I'm just task
Teresa Torres (51:20.295)
Yeah.
David J Bland (51:29.776)
transactional type interface for things. Because there's so much about product that is beyond beyond that. And I'm hoping it doesn't get kind of constrained into that little box.
Teresa Torres (51:32.999)
Yeah.
Teresa Torres (51:42.257)
Well, there are already software companies that will o do automated evals for you. And I'm gonna use behavioral analytics as an analogy again here. You know, you can buy Google Analytics and you don't have to buy it, it's free. Install it and you're gonna get some off the shelf analytics.
And every product person knows exactly how limited that is because you want to measure value tied to your product, which isn't a page view or a click on a link, right? It's a value creation moment that is unique to your company. I think that's a really good analogy for evals. Like yes, you can buy eval software that will give you off-the-shelf evals and you're gonna get really crude measures of quality.
And we really are going to have to define real evals that measure value creation moments that are unique to your product.
David J Bland (52:30.599)
Interesting. So qual and quant, it's still balancing regardless of what the tech is going forward for product. I just I just want to thank you for hanging out. I mean we could have easily just rift on things. Maybe maybe we can figure out a new venue for that. But I was just great catching up with you. I love how you show your work. I love that you embrace, you know, this world of AI with some skepticism, but also willing to see how it could make your work better and how you teach better. if folks
Teresa Torres (52:34.673)
Yep.
David J Bland (52:59.889)
probably know how to get in touch with you, but if they don't and they listen to this and they want to reach out, what's the best way for them to find you?
Teresa Torres (53:07.088)
Yeah, so I blog at Product Talk dot org. I should say I blog, I podcast, I host events, I do all the things these days at Product Talk.org. I am on LinkedIn. It's probably not the best place to reach out because LinkedIn provides terrible inbox management tools and I get slammed there. Honestly the best way to reach out is just send an email to contact at product talk dot org. and yeah.
David J Bland (53:36.261)
Awesome. So product talk.org, we'll put that link in the description and on the detail page. I want to thank you so much, Teresa, just for hanging out. I had thoroughly enjoyed our conversation. Hopefully people get a better perspective of how to do assumptions mapping and the different layers and levels to it and maybe not try to put people against each other. Also, not maybe not be so scared about AI and learn about hmm, maybe product and evals, that is something I should start learning. If not learning.
Teresa Torres (53:54.909)
Yeah.
David J Bland (54:03.065)
already. But I just wanna thank you for just always being open and honest and just awesome. And thanks for joining us for the conversation today.
Teresa Torres (54:09.843)
Thanks for having me. I really enjoyed it.