Transcript
Speaker: go out and preach the gospel of experimentation in terms of like, get people to just run experiments. Don't put guardrails up. A lot of times, like the reason people don't run these experiments is there's too much process.
Speaker: You've like, with good intentions created like, oh, we're going to make sure everyone writes these documents. We're going to make sure everyone fills out all the blanks and everything. It's a good practice to eventually get to.
Speaker: But if the current problem is no one's running experiments, doest those documents don't matter, right? That's a secondary thing. The first thing is getting people into the practice of running experiments, even if they, you know, do it kind of fly by the seat of their pants and just, you know, put it up there. I think that's still going to be more helpful for the business in terms of figuring out the direction things should go.
Speaker: Welcome to Unite Voices, hosted by Katie Green. Real stories from the people behind today's most innovative experimentation programs. No fluff, just wins, failures, and the lessons in between.
Speaker: Zach, thank you so much for joining us on Unite Voices. Hosted by me, Katie Green, I'm principal advocate at Chameleon. My favorite way to explain what I do is that I'm a community connector and I sell experimentation. I sell the practice of experimentation, the culture of experimentation, and I love connecting people who are solving really complex problems. One of those people who has immense experience in this is Zach, who's our guest today. Zach, you were nominated for an experimentation thought leadership award in the data and engineering industry.
Speaker: analytics innovator, i think, category. It's very hard to say all those words all at once. But you were nominated for ETLA. You have incredible experience with particularly numbers as a former background in some economy work. So can you introduce yourself and tell people who you are?
Speaker: Yeah, absolutely. I'm Zach. I'm an economist, background in economics. I worked for a bit at places like Amazon and Udemy and currently at FiveTrend. And I've done a lot of experimentation stuff, as well as a lot of pricing and other types of things. I think, yeah, that's kind of my background is I've done all sorts of anything that has data, numbers, and usually some connection economics. And I think that's why people are going to tune into this episode in particular is you, I mean, I myself have run tests. I'm an experimentation practitioner myself for the last 10 years or so. I've run pricing tests.
Speaker: I do not have the background in being able to measure that kind of, you know, experimentation rigor. When it comes to testing pricing and pricing models and subscription models and all these things, I relied heavily on our analytics folks to help me understand the impact of these tests. So I know a ton of people are going to be tuning in to hear your take because they're probably struggling with the same thing. They don't have access to someone, you know, as bright as you in economy and, ah you know, all of these things and in economics. But can you tell us a little bit about how that intersects with testing? Because of course, that is the nature of the show. But with your background in economics, how does your definition of a metric compare to that of a traditional product manager?
Speaker: Yeah, so I think I think i could say you know it's a good example of like how thinking, at least economically, helps choose metrics. right We want to pick metrics that will surprise us.
Speaker: We don't want to pick metrics that are like, you know we already know what's going to happen after you after you run the experiment, you know, your kind of classic example is like, say you replaced an element on a web page and you made it more prominent. You made it you made some element on the web page more prominent. If you choose your metric is like, how often are people going to click on that that element? Of course, they're going to click on it more often. That doesn't surprise you. That's not really what you're trying to learn. What you want to understand is the trade-offs. You want to find like, what is the thing that's going to be,
Speaker: what is the what is the part of this, what could go wrong, right? what What could go wrong and what could go right? And that's what I want to choose as my metric for the experiment. I want to be able to test which of those two things you know will happen.
Speaker: And kind of more of like a pricing or monetization type context, you could think about that as like promotions, right? If people run a promotion, it's not surprising that if that promotion comes with a discount that people buy more. That's not the that's not the thing that we're ultimately trying to learn. We want to learn you know does it drive whatever outcome we're we're trying to get to?
Speaker: um And we want to do that by picking metrics that like we we don't know the answer to which direction they'll go before before the experiment. So I think that's kind of like maybe how it intersects with how I try to think about it. um And you also, you know, you also just it helps to have like, you know, some statistical understanding of like that certain metrics move faster than other metrics. Usually these are metrics that are like closer to the actual intervention you're doing. So if you have a funnel and you're intervening at step three with the experiment, you want to look at metrics that happen obviously after that. But you want to look at metrics that happen at like, you know, step four rather than the ones that happen at step five, because you're going to get a lot more power and a lot more um impact.
Speaker: for from looking at those. I love that. I think power and impact are what people are looking for. and definitely the reason they're tuning into this episode today. i want to get into, i know you have completely reworked an experimentation program in your past. And that I think is going to be the big chunk that people are looking to do because maybe they're in a similar position. But I i think the context of your current role being it's a pricing and monetization scientist, right? Can you tell us a little bit about how you're working with metrics and data today and what this means for your understanding of the experimentation at large? And then I think we can get into a little bit more of the nitty gritty. I think one thing I've learned from this current experience is like to not be scared of running long experiments. Sometimes experiments take a long time ah to happen, but if the impact of them is large enough and it's a big enough you know business direction, you can't actually justify it.
Speaker: This is kind of you know an interesting context because it's very like B2B enterprise world. So there's like... um fewer you know fewer data points than you might have in, like say, like an e-commerce type type setting.
Speaker: um but So you you might need to run the experiment longer. But I think one of the one of the nice things there is that like you can can ultimately get insights that are are grounded in reality rather than um you know, the kind of vibes based decision making that you'll otherwise kind of like have have to do.
Speaker: um And so, yeah, I think that's kind of that's kind of the connection, I think, between experimentation and and a lot of pricing and monetization work. Obviously, like people usually aren't running experiments directly on price, you know, all they they can, I guess, in some context. But um you you're running you're running experiments a lot of times like on the on the periphery, right, you know, on stuff like i don't know how much of people get discounts for various things or or that, that kind of stuff. I think it's a, I like to call it the presentation of price. You know, it's, it's the emotional presentation of how much it costs.
Speaker: That's great. And when it comes to experimentation, I just know you have such a deep, robust experience here and you were with,
Speaker: on a program on completely restructuring it. Right. i think that's the project that I'd love for you to share with people. Now they understand the lens at which you're coming at it, you know, as somebody who does not have, I don't have a ton of experience in pricing, right. It is like the presentation of pricing, the emotions around pricing, um, as a marketing kind of experimenter.
Speaker: Can you tell us a little bit about the infrastructures that you were working within um, what changes you made to scale up experimentation. We have a ton of practitioners tuning in and that's their biggest question is how do I take it to the next step? So can you tell us more about that project that you worked with?
Speaker: This is back when I worked at at Udemy. And what i what I did there was like, first off, I knew nothing about experimentation before starting that job. I had never run an experiment or done any sort of experiment analysis. So I was kind of in that boat exactly. and but it But got a job to work in experimentation there.
Speaker: And the goal was to kind of like rewrite the entire like you know the experiment analysis, like how we do experiment analysis there. And one of the nice things about that was that I got to think about a lot of problems from first principles. So I got to think about a lot of, you know, I didn't have any sort of like preconceived ideas about how to do that correctly. So um what what I did kind of there was I think there's there's there's kind of three key pillars, I guess, to getting that right.
Speaker: which is, one, ah always doing experiments is better than not doing experiments. And we shouldn't you know like if you can get people to run experiments, that's a positive thing. Getting caught up on, like are they running it correctly, is like every I and dotted and T crossed is not really like a good way to go. We want to make sure that people look at some data about what happened when customers saw A and what happened when customers saw B. And the details are are good and important to get in a mature situation. But like the first part is just doing that, like getting getting to look at both those things.
Speaker: um the other The other kind of piece that I i would say is is important is to make sure that we're okay with like making decisions under, um not necessarily at this kind of statistical significance levels that are often used in like the academic world, like, you know, there's like 5% levels and and that kind of thing. I think the generally education,
Speaker: in a business context, you're willing to take more risk than some major government policy and in any it or some major government policy or like drugs you're shipping or whatever, airplane design. like Here we're talking about like you know e-commerce or something. right like it not no one's no one No one crashes if you get this wrong. So like the kind of the key, the key thing to do with there is I think to choose significance level that makes sense for the industry, for like the industry and the projects you're working on.
Speaker: um And ah and the the reason for that is kind of twofold. One is that like, you want people to first off look at statistical significance and if they always get it um ah if they always get insignificant, right they won't look at it anymore and they won't think about the he the size of the air bounce. So you want to be make sure that you kind of like hone in on that and get get some kind of thing where people actually see it often enough in your data in a reasonable timeframe.
Speaker: um like you know before before they could have to make decisions um the other thing is like is kind of touched on a little earlier was metric selection uh metric selection is i think one of the key things um standard reflex of people ultimately everyone's trying to drive revenue right so like everyone is trying to drive revenue so their standard reflex will be to pick metrics like revenue um because that's you know if revenue goes up that's good uh we all we all know this um the problem is that everything in the entire world affects revenue right macroeconomics people's demand for like you know that that particular product class your competitors entering something the social media intern tweeting something bad whatever like everything is going to funnel in to that to that kind of like revenue decision
Speaker: And so the and so, you know, given enough time, your experiment will be able to distinguish between like whether your intervention increased revenue or not. But it'll take a really long time to get to an answer there.
Speaker: A better way to do it is to pick revenue is to pick metrics that are closer to what your actual, um you know, like closer to what your actual and innovation is. So, you know, if you if you improve some part of the funnel, you know, look closely to it and try to find metrics that have you know significant power and that's kind of like the other step there which is getting people to run power analysis um and look at uh what can happen um look at like realistic sample sizes that they'll need for their experiments the other thing i always like to encourage is like um when you're trying to do like peaking discussions of like how often should people look at their their experiment results is like
Speaker: the The correct answer isn't zero peaking, and the and it's not like peaking every day. it's it's something It's something in the middle. and the way you want to structure that is with pre-specifying how many times you're going to peak um during the experiment and then calculating like the correct statistical adjustments for that when you're making when you're making like decisions.
Speaker: um You want to get to a world, like you won't get there initially, but you want to get to a world where people are mostly deciding to ship things that are statistically significant and not ship things that aren't. um And i that that world takes a while to get to, but and it's you know there's obviously practical things that come up. But the the the you know that's that's kind of the end state goal I think you you want to get to when running a program.
Speaker: I'm guilty of peaking. My bad. my bad. i mean, i I would do it daily. Because it's mostly that the thing that I struggled with when I'm running tests is what percentage...
Speaker: you know, decrease in a KPI? Am I comfortable with before being like, oh, no, this like really isn't working. And I know you have to wait until it has statistical significance, obviously. But it's like, okay, well, it's like a, you know, we have 100,000 people in this test, and the conversion rate is down 60%.
Speaker: sixty percent You know, that's obviously really dramatic. But like, having some level of statistical rigor to understand when it is premature, what what that Breakpoint is, you know, is something that I've really struggled with in my career. um So I hope that somebody listening to this is going to be like, oh, I have questions for Zach. I'm going to DM you on LinkedIn, try to understand where to go with this. But I think that something that experimentation does, and you clearly have a lot of experience in this, is it resolves...
Speaker: questions, right? That's kind of what we're here to do is we ask questions, we provide answers to them. But what that kind of creates in a corporate environment and probably particularly in B2B is maybe conflict isn't the right word, but you know friction makes fire. We'll call it conflict. there's The philosophy clashes when you have some level of data evidence.
Speaker: And so I know a lot of people experimentation struggle with that. I can tell you myself, I've had calls with CMOs and CEOs and I'm like, hey, this like idea you were certain was going to work.
Speaker: We tested it every which way and it super doesn't. Right. So can you tell us a little bit more? You have this leadership position. You've done so well in creating rigor and ah helping build this culture of learning and failing forward.
Speaker: I just would love to hear about how you handle those deep organizational conflicts when it comes to evidence that you're presenting to counteract a philosophy that may or may not be in place. But, you know, tell us a little bit more about the experience you have there.
Speaker: Yeah, absolutely. So I'd say like the the first thing is one of one of the powerful things about experimentation is that it kind of gives you the most ammo to say something that isn't isn't working. Right. If you if you kind of think it's not going to work, like say some executives idea isn't going to work because you've you know you've seen some like descriptive statistics or some kind of observational analysis, you know there's going to be a lot of conversation back and forth about whether you checked that thing or whether ah that assumption really holds and and all that kind of stuff. One of the like the good things about experimentation is that it is kind of like, look, we showed half the people this, half the people that, and the people we showed that new thing did not do so well. um
Speaker: So you know I think it's actually very helpful in forming those conversations. um the, you know, sometimes when it when it's when people are not convinced by that evidence, that the truth is something like, you know, that there's there's something besides that, and that innovation they have in mind. So it's often good to like try to extract, like, what exactly, like, why exactly does this not convince them, as usually the first thing to ask?
Speaker: um Because, ah so you know, there's, is there's this kind of thing where like, you're at like, a local optimization versus global optimum versus global optimum. And they want to move the business in a certain direction. And they understand they're going to take some hit initially until you start you know optimizing that new thing. You can think about like, so you've got some like e-commerce site, you're going like standard experimentation setup and you've like done it um you've you've highly optimized that site and we decided now we're going to take a 180 turn in sort of our like whole pricing model like we're you know we're no longer going to be like a la carte like transactional site we're now a subscription site something something dramatic like that like you know that that's going to initially hurt revenue because you haven't um
Speaker: I mean, you don't know, that but it could initially hurt revenue a lot. And it's a, you know, you haven't optimized for that. You haven't built anything like that. And, you know, so the the leader may be just saying like, we're making a pivot. Like we think this is, you know, like a a structural thing we're going to do. And now the question is just all about, like given that we're doing this, let's try to make it good.
Speaker: um and so I think that's important to understand the like the context of the decision of of of the kind of like concern about the experiments. um But in general, I think it's one of the best ways to like convince people. like Most people I've found are convinced by like experiments. like if if if they If they beforehand thought like something would work and then you ran an experiment on it and it didn't work, like in general, my general experience has been people then, you know, maybe pivoting to something else. Like but they generally will agree that this is, you know, not not convincing people.
Speaker: you know it's convincing ah that they shouldn't we shouldn't go forward with at least that version of of the idea. they may they you know A lot of times, they'll come back with some like iteration. right like you should i still want to I still like this like general framework, and I want to try some ah some different you know variation on that theme. But experimentation is a very convincing way to talk with docker folks.
Speaker: Sometimes I think our executives are executives listening to this, you know, I'm like, I feel like all the time I'm like, so the hippo's opinion, the hippo, right. The highest paid person in the room's opinion um is like something I talk about a lot. So I'm more like, Oh no, we're executives listening to this and thinking, Oh, like this is what the people talk about. No, it is. It is very interesting. I think it is but how I, I feel like I have a similar story to a lot of people coming into experimentation where we fell into it because they're we had a question and we were like, is this right? And I remember my first test, it was actually an email test. It was ah really not website experimentation. I was like working life cycle and multi-channel testing through paid. And i remember being like, how do we know this is going to work? And then we're like, oh yeah, let's AB test it. Like, I remember that being the first time I had that thought. So being able to have that conversation with evidence is, is really important. And
Speaker: There's one other thing that you mentioned earlier that I wanted to make sure I underlined that I didn't get to in that last question. But the question before you mentioned something that actually came up at our Unite conference in London recently.
Speaker: Earlier this month, I had the pleasure of emceeing our sold out ah Unite Summit London. It was very fun. i i had a really good time in London. it was very fun. And one of the things that came up in the panel was directly, it was kind of changing how you're thinking about your metrics from test specific metrics to that global change that you're talking about. How are these metrics affecting your business, your business metrics, your business KPIs? I feel like so many people do decision trees and we all understand that our metrics have to ladder up into that, but it's really hard to have a conversation about decision-making and, you know This corporate conflict and having that evidence kind of means nothing if the evidence isn't speaking the language of the person making the decision, right? So i I love how you're saying this. I think it's extremely important. And yeah, if you're not speaking the language of the decision maker, you have a really hard time.
Speaker: Proving any point you have with evidence. Before we get to our last question, i think I did – I do just have one more question for you because I i think you have – before we get into the final question, which is my Monday morning advice, which I'll explain later, i have one more question for you that I'm realizing I do want to cover before we – cut off the recording, which is marketing versus engineering, right? Is kind of a big question versus data versus design, right? There's all these different teams that touch experimentation. And because you have this incredible experience, like rebuilding an experimentation program, how do you create an evaluation criteria
Speaker: that both marketers and to developers and analysts, right? How are you helping lift up all of these different teams that are going to, you know, the ultimate decision makers and saying, hey, your philosophy is wrong and here's why.
Speaker: How are you managing the metrics with that and, you know, balancing the different teams approaches to experimentation? Yeah, I mean, I think, yeah, there's definitely different if different needs, I guess, in those in those teams. Like, you know, marketing it's usually needs something that's that's fairly easy to deploy and that has the ability to kind of like, they run a lot of experiments, right? Marketing tends to run like,
Speaker: ah a much larger number than people who touch, say, like the back end of the website whenever they launch an experiment. you know they You know, they might have but you you know different variations of an email, different variations of, you know, kind like text message or whatever they're doing for for advertising or push notifications, that kind of thing. I think ah so. So they're going to want like in a platform that's and a platform and like experience, it's like they can run a lot of things they're going to want to be able to like template size right like you want to get them to the point where they can like each experiment has kind of like here's the metric here's you know they look alike there's just one that has like a metric that you're going to make a clear decision on um and the other thing about marketing is a lot of the experiments are kind of like um
Speaker: I wouldn't say they're even like necessarily decision drivers, right? Like you're, you ran it once and now you're going to run, you're kind of to to figure out like what works better as like a general principle, but the next marketing copy is probably going to be different than like the previous marketing copy a lot of the times.
Speaker: Uh, so, you know, it's kind of, it's kind of like one of these things where you have a lot more these kind of quicker, um I need to gather information. Usually they don't last very long. And then you've got something like, like a product side experiment where the You need the ability to like also change stuff on the back end and like be really integrated into the site, which is also something that's going to matter for engineering and and for those teams.
Speaker: um And so you want to go out an easy way to to do that. And and you want a way that like you want to make sure that everything's instrumented. that's another kind of This is kind of one of those technical things about experimentation in like, you know the tech world is you want to have lots of you need to make sure these metrics are actually exist and like you can measure them.
Speaker: um And that actually takes usually a ah lot more time than and people think, because when the product starting, people are just trying to, like, get the stuff up, make the business work, make some kind of revenue and like exist. The kind of like the details of like making sure this like button click like triggers something in a database like is not usually like the first kind of thought in mind. So like one of the first steps of like getting onto a more experimentation forward thing is just getting those metrics to instrument it and existing.
Speaker: And then that makes and engineering's life much easier going forward. is like it's a fixed It's a fixed cost. You pay once, and then you kind of you kind of you don't have to pay it as much. At least you might have more metrics over time, but you don't pay it as much um once you've paid it once.
Speaker: And I think so i think that's that's kind of a key thing for like engineering. Products are run longer experiments, but they're going to have a wider variety of metrics. right like The metrics are often going to be la carte to like the specific experiment they're running about their particular product, and maybe it's a completely brand new product. So you need the ability to like, um do do kind of very flexible things on the back end. It's sort of the way to kind of like satisfy all these different groups is yeah making sure you've got like a setup where you can, you know, you can both run
Speaker: fast experiments and you can run longer experiments and you can brain change the back end and you can change the front end and whatever the mailers are. You want something that kind of like encompasses all those surfaces.
Speaker: um that's that And the the key thing is encompasses all those surfaces, but at its core has some kind of unifying understanding of like this person is mapped to variant one this person is mapped to variant two like and there's kind of like a similar way that you that all this data gets structured and comes out so that the analysis becomes like automatable straightforward and um easy to easy to understand and make decisions from um A key thing, kind of last thought on that that sort of front is like um that one of the key things is to make sure you you commit you decide how you're going to decide before the experiment starts.
Speaker: Otherwise, you run into this kind of like ex post thing after the experiment's done and everyone kind of looks at the metrics and squints and cuts them every which way and reads the tea leaves and comes up with some kind of story that kind of usually convinces you of your priors.
Speaker: um And so the i think like the really like the very powerful thing is to say, like before the experiment, if this metric goes up, this metric doesn't go down, whatever, then we're going to launch it.
Speaker: And, you know, kind of like decide beforehand what what you're going to do. And that's a good way to get everyone on the same page because then you could refer it after the fact, hey, look, here was our decision criteria. And, you know, why why would why what have we learned that would make us change that? That's a good way to kind of organize everyone around that.
Speaker: I think that's the most important piece that I want to underline, write an article about is, what is our decision-making criteria? And you have to like provide a reason why we're divorcing from that if so. um But having that in advance, I think also the like flexibility point that you made before that really critical is no experimentation is one size.
Speaker: experimentation program is one size fits all. And you do, you know, you have a really good breadth of experience in speaking to that. So thank you for sharing your knowledge on that. And with our final few minutes, I ask our guests all the same thing before I let them go, which I call the Monday morning advice. I realize it currently is Monday morning for us. This episode is going to air on a Thursday,
Speaker: but it's It's the next day. It's the takeaway. What are people doing? What is one thing you can recommend people do? And so I want to speak to your experience of, okay, you're maybe, I'm trying to give you a little bit of a scenario here, but maybe someone is building a new, brand new experimentation program or their experimentation program has stalled and they're looking to scale up and make some dramatic change.
Speaker: You know, what is the one thing you think somebody should start with to do that? Yeah, so I think the i think the one thing that you should definitely start with for getting those programs play up or expanded is go out and preach the gospel of experimentation in terms of like get people to just run experiments.
Speaker: Don't put guardrails up. the they kind of a lot of times like the reason people don't run these experiments is there's too much process. You've, like with good intentions, created like, oh, we're going to make sure everyone writes these documents. We're going to make sure everyone fills out all the blanks and everything.
Speaker: And it's like a good it's a good practice to eventually get to. But if the current problem is no one's running experiments, he doesn it those documents don't matter. right that's That's a secondary thing. The first thing is getting people into the practice of of running experiments, even if they you know do it, kind of fly by the seat of their pants and just you know put it up there. I think that's still going to be more helpful for the business um in terms of figuring out the direction things should go.
Speaker: So that would be kind of my my first recommendation would be to like look at what guardrails are currently in place or what I mean, guardrails probably, but what like ah what processes are currently in place that people have to do that create friction and just take them away, even if they're good, even if they have good sound reasoning behind them. If the problem is no one's running experiments, just take away those things, get people to run experiments and then try to teach best practices um while they have actual experiments they're looking at.
Speaker: um That just becomes a much easier conversation than you know, that when there's no experiments and you're trying to teach best practices about something they don't currently do. So that's ah that's kind of what my thought process there.
Speaker: I love that. And we'll clip that every which way for sure. It's kind of practice what you preach is what you're saying ultimately is what I'm hearing is, okay, if you're preaching that you want to do experimentation, you want to fail forward, you want to have a culture of learning, what you have to do is kind of look at yourself, have a hard look in the mirror and say, are these processes actually pushing that culture forward? So I i think that's a wonderful piece of advice to end on.
Speaker: Zach, thank you so much for being a part of Unite Voices.






