Zencastr
00:00:00
00:00:01
Speed1x
Format▸
Share
Embed
Report

Why do people ignore good A/B test results? A statistician's perspective | Cristina McGuire, Experimentation Data Scientist at Microsoft

Unite Voices
Unite Voices

8 plays · Sep 24, 2026

Transcript

Speaker: <unk>s kind of one of like the learnings that I have throughout being in this space of having a really like strong statistics background and really like more layering in the business and empathy you know like of like how people are doing it. With the statistics background, though, what I really like about it is like it's easy to make some of the conversations less emotional. So then, say, for example, you you talk to a customer and like, hey, we really want to launch this experiment, what we should do.

Speaker: And you can kind of like bring back to kind of like a neutral zone by like providing them like, well, did you power your experiment? So like you kind of go back to like the structure, because if you have the structure, when they talk to like their leaders, it's not about them. It's just the math and it's just the data. So I kind of layer in like the statistical rigor. i think having like that conversation and understanding like the business need is what really makes bigger impact.

Speaker: Welcome to Unite Voices, hosted by Katie Green. Real stories from the people behind today's most innovative experimentation programs. No fluff, just wins, failures, and the lessons in between.

Speaker: Welcome to another episode of Unite Voices hosted by me, Katie Green. i am serving as the principal advocate at Chameleon. And I am joined by Christina, who has incredible experience in experimentation, has built programs at companies like Chewy and Expedia, and is most recently joining the team to do experimentation at Microsoft. So Christina, can you introduce yourself and tell the people who you are?

Speaker: yeah Yeah, I'm Christina. um I entered the world of experimentation more than five years ago as a data scientist. So as Katie said, I'm currently working with the Windows experimentation team at Microsoft. Before that, at Chewy, and then where it all started really as Expedia Group, I just came in as a data scientist, put in a place where in The main goal was to drive experimentation, and I've never turned back ever since. So I really love being in this space. I see it as a niche, like maybe like three years ago, but you all know like the most recent time, like experimentation this feels like more like a bigger community now, and I'm really excited to be part of it. And I feel like because it is that we're a tight knit group, right? So we do podcasts like this and go to events and whatever it is. And it just feels like it's getting bigger. And I'm so your name has you and I were connected from a a connection of a connection, you know, so it's amazing just showing in real time how people can get connected in this community. And that's one of the takeaways from the podcast that I really want is somebody listening to you and saying like, oh, I really want to follow you, connect with you on LinkedIn. That's success in my eyes. So for anybody who wants to connect with Christina, please do, or myself, we're always here for you. But I mean, you just name dropped huge brands.

Speaker: So tell us a little bit about that. I know a lot of people listening probably work at smaller to medium sized businesses or businesses. They're working on smaller teams, maybe not necessarily smaller brands. I know a lot of big brands use small teams. So can you tell us a little bit about what it's like to work at scale? I mean, everywhere you've worked has just been, i mean, huge experimentation programs.

Speaker: Tell us a little bit about that. Yeah. um So having a background in statistics, I started my career just doing general data science. But what really got me fascinated with experimentation is that um I joined the team and then I realized the most practical application of stat statistics can be scaled in real world. Like technically you're running T-tests like hundreds of times, like thousands or even more, right? And just to get started with that, I thought that's like, wow, this is really um interesting. And as what you mentioned, I became really like deeply involved in like experimentation community, attending conferences, networking events, and learning learning from like practitioners of experimentation. um i think like with the scale and how it works out is that I was lucky when I started my my experimentation journey. i came I joined Expedia Group where they already have a mature kind of experimentation, but they were rebuilding to kind of like make things like even better. So it still felt like we were trying to

Speaker: um build the experimentation culture but now powered by a platform and I think like having a really good platform is really what makes what makes experimentation to scale even better because um as sta is a statistician we know like the right thing to do right like this is how you this is how You should run experiments like as if like you're a scientist, but you cannot tell people and watch every single person um and check them if they're running really good experimentation. So working hand in hand with the platform is really what I think drives... um

Speaker: the scale that would make your experimentation culture really strong. um Actually, like when I moved from Expedia to Chewy, what got me interested with Chewy is like they have a little bit smaller experimentation program. And I was hoping to build like the to bring in the foundations that I got from like more mature companies like Expedia and kind of like bring that to life like from scratch and like building the foundation of it so that like people can run more experiments do it it with the statistical rigor and the level that we're in they could really use it for decision making and we know like you know the Microsoft is another big company each organization have different like culture and the way to scale is also very different

Speaker: yeah No, that's wonderful to hear you Your experience is so rich. And I know a lot of people are going to listen to this because of that experience. And ah I'm curious. I didn't prep you this question, so I'm sorry it's coming out of nowhere. I'm just i'm very, very curious to know.

Speaker: You have a master's in statistics. A lot of experimenters do not have such rigor on their teams, themselves, whatever it is. And I'm curious, how do you teach this? Like when you're working with teams, like you're saying with Expedia Group and in specifics, how are you working with people to show them how to have that rigor without being a math teacher? Maybe you are a math teacher.

Speaker: I did teach math for undergrads as part of my teaching assistantship, but like that's like my my my my experience of that. But that's a good question. And I think like i'm maybe I'm going to go to a different route in answering question, but I'm going to answer it and see where where're this conversation lands. So I think like when I started knowing maybe from the my manager that they have statistics background and also within my team, like I've worked with people with PhD in statistics and, um you know, like really like um scientifically driven people. um So when I started my my career at Expedia, ah my role was to that was given to me was to raise the bar of experimentation. So when I when I and I took it seriously, right, because like, you know, like you're coming. I've already been few years out of school, but but still like you carry that like this is what I learned and this is how people should run experiments. Right. So I went in with the mindset and thinking that like, OK, my goal is to identify the ways how teams are experimenting and maybe identify what's wrong and maybe I can teach them the right way.

Speaker: That didn't last so long because like I realized that like everybody wants to run good experiments. It's just that they are dealt with different limitations, constraints. Sometimes they have limited time, incomplete data. There's organization pressure to prove that you're bringing in conversion or like the product is not working, but then the business really wants to push it. So there's a lot more. in that goes into experimentation that really like statistics can help um help a little bit but like it's not the full story so kind of like that's kind of one of like the learnings that I have throughout being in this like space of like having a really like strong statistics background and really like more layering in the business and Erika Endrijonas

Speaker: are doing it but with the statistics background though what i really like about it is like it's easy to make some of the conversations less emotional, right? So then say, for example, you you talk to a customer and like, hey, we really want to launch this experiment, what we should do. And you can kind of like bring back to kind of like a neutral zone by like providing them like, well, you know, like um did you power your experiment? So like you kind of go back to like the structure because if you have the structure, when they talk to like their leaders, it's not about them. It's not about, it's just the math and it's just the data. So I think like having that,

Speaker: That's kind of like the way I kind of thought or I kind of layer in like the statistical rigor. But you cannot just be like telling people like, oh, you have to like do power analysis or.

Speaker: your hypothesis is wrong. It's really like, I think having like that conversation and understanding like the business need is what really makes bigger impact. No, that's a very helpful answer to that question. And i want to kind of flip it around actually and ask you the question in a different way. It's a different question, but it's kind of the things your you your answer got to, which is so much of testing in my experience is things that are almost immeasurable. How do we measure brand impact?

Speaker: You have to create the right hypotheses and track the right metrics. What experience does a statistician have in that part of the workflow that is kind of taking this amorphous thing and saying, oh, I need to turn this into black and white numbers for the thing that you're saying, which is I'm here to help you make good decisions, make easy decisions. What is it like translating your experience from kind of going from the emotion to the practical?

Speaker: yeah Yeah, I think like when you have like the statistical background, it's very structured, right? There's like a path like, okay, let's design the experiment. um I understand that the goal is to identify or measure like brand. What data sets are available that could be a proxy of like how well like the customer likes your brand?

Speaker: do we have Do we have that data? So I think like it always starts with, I know it's a cliche, but it's always starts with having a good hypothesis, understanding kind of like what measures um are measurable, what metrics are measurable, helping the customer ask like the customer in terms of like, okay, can we quantify brand, you know, like in terms of like,

Speaker: um In the perspective of the business, the revenue impact, in the perspective of the customer, can we understand maybe like the retention rate or like of the product? Like how can we measure like the the goodness of your idea to like the product itself? Is it like engagement? So I think like um a lot of the time is like grounding them back to like what's measurable and what's what can we answer, right? Because like if we have all the power in the world, we would want to know the answer to like the exact thing that we want. Like, you know, what is the impact of like this feature to a brand um long term? But we know there's a lot of like limitations with experimentation wherein you only have like two weeks, maybe four weeks to to measure. So maybe we wouldn't get that. And you can go through different, like more advanced statistics to kind of solve that problem of like long-term metrics. But is there a simple way wherein we could go back to the basic of like experimentation, of designing a good hypothesis, identifying the right metrics, um understanding if like engineering can like launch ah an experiment that you want. And like what decisions can we make from from this insight? So at the end of the day, you want to maximize the learning with with the reasonable resources that you want to put in, right? So I think to answer that question is like, as a statistician or as a data scientist, like you kind of ground them back to like, what's measurable? How how can we implement this? and What type of decisions are we able to make? And is that good enough? If it's not good enough, let's go talk about it more

Speaker: And I think a lot of people are going to resonate with that. I think it sounds, I know we say it's like oh, it's so obvious. Start with a good hypothesis. But there are so many people running so quick and their expectations are so high and the stress is really high that they're like, ah, I got to get these tests out. And it doesn't really matter what they are as long as they're out. um You know, it's it's almost, I did see a really good presentation at a conference once that was saying they don't measure by how many tests they run. They measure by how many decisions they make.

Speaker: based on data. And I'm like, wow, that's, that's actually really good. Cause then it really forces you to focus on like learn rate, you know, and there's, we could go down that rabbit hole because there's issues with it as well. But like, it's all starting with a good hypothesis. It sounds, you know, elementary, but it is so easy to lose because it is such an elementary critical part of the funnel. So yeah. I think like sometimes in the life cycle of an experimentation, when a person asks the question in the middle, like say,

Speaker: maybe we'll get through this question that like should I run the experiment longer like you wouldn't even need to go through that question if you just had like spent time earlier to kind of like develop your your um your plan right and I have maybe I'm biased but I think like most of the time you only have to develop this plan strategy like maybe once every like month and reassess it because like it's replicable that's the fun thing about experimentation is like you run one good experiment, you get good at it, and then you can just run your experiment over and over again with different features. So so yeah, um maybe we're simplifying it down, but I think the idea, a lot of the people a lot of people get overwhelmed with like the details, but then if you just expend time strategizing on your experimentation program, you could scale it, yeah.

Speaker: I love that. I obviously wrote that down as a really good note because it is it is replicable and that's that's part of what makes it fun. i think something I personally struggle with, and I'm kind of selfishly like asking you this question, what I've run into in the past is it's kind of easy to find an insight. It's like looking for a needle in a stack of needles where you have so much data, you have hundreds of metrics available to your testing program. I just want to get a little bit more tangible here because I'm very curious about your perspective on it.

Speaker: How do you, is it that first strategic piece that you're like, okay, this is the only decision making metric we're using for this test. Do you end up flexing into other decision making metrics based on the performance? Can you tell us a little bit about how you prioritize the hierarchy of your metrics and what that does to the impact of the test?

Speaker: Yeah, this is where probably my statistician kind of mindset would still kick in that like, I think like sticking to like your hypothesis in terms of like identifying what was your primary metric, secondary metric in guardrails, right? So you you have like cards to kind of fill in before before you start the experiment, so you kind of just need to use it wisely. But I think, like, having, like, a good primary metric is really important to make sure that, like, you are sizing your test, like, well enough. Like, if you, for example, like, in the middle of the experiment, like, oh, that one looks good. Let's make that our primary metric. But then, like, um maybe you're only observing it because it's, like, you know, short. You didn't really have enough power to, like, ensure that the change that you're observing is, like, not exaggerated, right? So I think, like,

Speaker: um always kind of anchoring towards like one primary metric or two is really important um to kind of again to remove a lot of this like decision analysis paralysis once you have the data because there will be data aside from the metrics that you have listed right so um i remember back in the day um depending on which platform you're looking at if you're like a purist like statistical purist like have as less metrics possible because that would reduce like your multiple hypothesis errors, um your false positive rate will be lower if you just like become very intentional. But then there's another side of like, we live in like the two, like, you know, in the generation where data is really just like available. So like might as well give all all the things, all the data to the user. So there's like pros and cons with that, right? Like when you have all the data, most likely some of them will just be randomness that you're observing.

Speaker: But then if you become too intentional, wherein you limit your data to only, like, say, 10 metrics, the the the the audience or, like, the the experimenters feel like they don't have enough, right? So I think, like, how do you balance that? I think, like, um what I've seen work working is, like, having...

Speaker: allowing people to add experiments and it's really like going back to your engineering capabilities of how much metric can you calculate but like if you have that that capability enable a lot of metrics but then make sure that users are specifying what their real primary metrics are and like secondary and card drills and also like educate people that the the primary, secondary, and guard drills are the things that have their, they use, you use it to make launch decision, to make a decision. And then all the other metrics are there available for you to maybe generate the new hypothesis. Maybe there's a new insight that you didn't look into. So that there's value in both like a set of metrics that you're going to use to make the launch decision and then value of the metrics to kind of like

Speaker: further enrich your learning about the experiment or the feature that you're testing. Because, um yeah, like, they they have different use cases and you can definitely, like, leverage in both. But just always think about what will make it easy for you to make decisions or are you just going to confuse yourself in the end? Yeah. I mean, i confuse yourself in the end is is probably what I've done a lot in my career. If I'm being honest, that's why i was like, selfishly, I'm going ask you this question because it's something that I struggle have struggled with in the past. So as I continue my career, don't be surprised if you get a DM from me.

Speaker: Being like, can you just help me with this really quick problem I'm having? Well, I think, you know, we've done a really good job of talking about exactly how to be rigorous. And you've kind of briefly mentioned, you know, it, it kind of takes the emotion out of it. But what do you do when you run into, like,

Speaker: people who maybe don't want to accept a reality, right? This is this is the test outcome. This is what it's sharing. But somebody is trying to poke holes in it or that because I've seen that a lot myself, too, as you're working with people and they maybe really want a feature to go live and we're showing that conversions are down. But next step clicks are up.

Speaker: But that's not really that's not what the hypothesis was. Like, how do you navigate a situation like that where somebody doesn't really care about the math? Yeah, yeah, that's a good question. And definitely i've seen times like that, like that's where you kind of need to press the brake of like not freak out, like why are you not seeing this, you know, and kind of like put yourselves in the other person's shoes and understand like what they, what, again, what stakeholders they have to deal with, what organization pressure that they have to deal with. And most of the time i would suggest to like um kind of like agree with them, like, okay, like I see what you're saying, but we should also like quantify the trade-off, right, to kind of, like, say, like, hey, we're seeing engagement to be, like, really high, but actually, we're not seeing, like, enough change in this, like, primary metric, so I think, like, just making sure that the experimenter or the user is, like, presenting the whole data and not just, like, filtering them, right, so I think i think my way of doing it is to just

Speaker: yes but or yes and yes and let's show like this full story um and i think again we we harp on it as like if if you have like a good hypothesis and like the metrics you always have to be presenting those no matter what the other insights you have may have gotten right so um that's one way um i think like you could go deep. It depends depends on, like, which customer you're working on and what problem that they're trying to solve. Like, say, if the problem is, like, oh, there's this metric and it's improving, the question is, like, okay, um can you convert that to, like, a revenue? Like, if, you know, like, sometimes they're saying, like, oh, the click-through rate in this specific button and button is really high. Should we launch this? And I'm, like, okay, if you how much dollar does it mean for, like, more people to click that, right? So I think sometimes, like,

Speaker: the the kind of rule of thumb whenever you, when you do like site opportunity sizing is like, put it I know it's not it's not easy all the time, but like to put like a dollar value against your metric and see like, is it really important? Is it like what they call practically significant? this Will the business really care And how how does that look like in the trade-off of like how much effort you're going to put into to the cost of like launching this feature? So yeah, kind of like battle it with another data point, I guess. is it No, I think that's perfect. We actually at Unite... the summit london our conference in london there was we i led a panel with a few folks and one of the main points was every single metric on your experiments should ladder up to a business objective and that's kind of ah another way of saying it is making sure okay what is the actual value of a click through does it outsize this decrease in conversion like that that could be the case who knows like

Speaker: I've seen crazier shit when it comes to experimentation. So I love that. But I just really thought it was important to underline the stick to your stick to your point, like stick to your this is our primary metric because this is our hypothesis. And you've done to your very first point to bring the summary all the way back to the beginning. You've done the work and.

Speaker: had the discussions of making sure your strategy and your hypothesis is really good. And when you continue that work into your metrics, it kind of filters into the decision at the end where you're able to be strong in what you're presenting and say, I know that this is the right piece. So for everybody listening, there's your quick little recap. I realize we're running out of time quickly, but there is one question I typically like to ask at the end, which is what I call the Monday morning advice. um I should really trademark that. But it's really just... is somebody listening to this on a Friday? Probably. I hope so because our episodes release on Thursdays. So I'm, I mean, maybe they're listening to this on Thursday. If so, you're my favorite listener ever, but like the Monday morning advice is just somebody who wants to get better at making their experiments more reliable. They don't have a statistics background where, what can they do to level up and try to implement some of this rigor and strategy that you have done such a good job of maintaining throughout your career?

Speaker: Yeah, um I'll answer it two parts. One, in the perspective of a data scientist, because I think like I've gained a lot of, like I've shifted my mindset about experimentation, and I'll also like answer it it if you're just an experimenter. right So I think like my advice for like a Monday morning, if you're either working in the experimentation platform or you're a data scientist like working ah embedded in like the product team. I think like my advice is always to understand the problem of the customer. And that doesn't mean it's the question that they ask you most of the time.

Speaker: Because most of the time, if they ask you, is this the right MDE? And then you start calculating like the MDE. yeah Actually, like the the question that they're trying to ask is mainly about like um how how long do we need to run the test? like What's the cost? like you know There's like more into it. So I think like over time, as like a data scientist in in the experimentation, I think most of the time, i see us like jump straight to, OK, what is the metric? and What method can we apply? What is the data design? But really like understanding like where the question is coming from. Sometimes you need to zoom out really um to to really understand that is like um um really important. um And I think um also recognizing that the barrier could be different, um could be is rarely the same across different teams. Sometimes you know people are trying to, they need to improve like their experimentation knowledge. Sometimes there's really problem in engineering and there's really not much science to the same thing with data quality, right? So I think, um yeah, like any any experimentation problem, because it's such a cross-functional field that um you really need to figure out first what the problem and what the customer needs is before you jump in into the conclusion. um And then on on a regular experimenter, like, good news is that you don't need to know all the stats. If you have a good experimentation platform and and that you trust, I think, like, the rigor and the quality of experimentation has just gotten higher and higher over the years, especially with platforms like Apple and Statsig, like, interesting. driving the scale. a lot of i

Speaker: I've done a lot of like um benchmarking of different like experimentation platforms, and the basics are there. And I think like as a as a starter, don't get overwhelmed that you need to know the stats because most of the time you don't. You just need to trust that. like um ah Find like the knowledge repository of the platform that you're using. Understand what's the step.

Speaker: And there's a generic step of like how you should run the experiment. Do it one time. and then do it over and over again until you become the expert. Most of the people in the experimentation world that are experts in the domain that they're in, they probably only have to run like one or two experiments for them to be experts because a lot of the people just shy away from it or they they just just don't see themselves as like, you know I need more stats, I need more engineering to like run an experiment. But really, you just need one or two times to like really go through the process and understand And then you'll gain so much knowledge and hopefully you'll be able to see experimentation more as like a tool for you to make big, better decisions aside and not kind of like, a ah you know, like a process that delays kind of your product development. So, so yeah, you don't need to know the stats. That's kind of the advice. Just like go with it. leverage your your data scientists in the team and and they'll help you. I love that. It's like, I feel like that's not really the answer that we typically get from somebody who's like, you don't really need me. But no, I love that. I think it's something everyone can take away. But it was wonderful talking to you I realize we're at time here. So thank you so much for being on Unite Voices, Christina.

Speaker: This was so fun. Thanks, Katie.

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Recommended