Consensus on Methodology and Metrics
00:00:00
Speaker
What we can all agree on to start with is on the methodology as well as metrics that should be used across the board. Because for a delivery space, we're a marketplace across consumers, dashers and couriers, and merchants. And we all agree that for consumers, we want to have the the best quality, reliability, satisfaction of orders that they're getting, not just focusing on short-term, but rather on the But then we allow local you know brands to to actually go and iterate on those common common metrics. ah But then only and now, after kind a few years, we're actually consolidating and building a unified platform, but always starting kind of different foundations first.
Introduction of Host and Guests
00:00:48
Speaker
Welcome to Unite Voices, hosted by Katie Green. Real stories from the people behind today's most innovative experimentation programs. No fluff, just wins, failures, and the lessons in between.
00:01:03
Speaker
Welcome to Unite Voices featuring Ilya, Makram, and me, Katie Green, your host. I'm the principal advocate at Chameleon, which means I am responsible for our online community and creating a space for experimenters to learn from each other. i am joined by Makram and Ilya, and I'll let you both introduce yourself. But Makram, why don't you go first?
00:01:24
Speaker
Hi, thank you. Thank you for hosting me. Yes, my name is Makram.
Makram's Introduction and Background
00:01:28
Speaker
I'm head of Marketplace right now at IDMeat. And before that, I was managing experimentation at Intuit, Intuit experimentation platform. And you're going to see a lot of common things between me and Ilya on that. And before that, I was at LinkedIn managing T-Rex, LinkedIn's experimentation platform.
00:01:47
Speaker
Very passionate about experimentation, personalization. So we're excited about this session today. And Ilya, tell us who you are.
Ilya's Experience in Experimentation Platforms
00:01:56
Speaker
Yeah, definitely. Ilya, thank you, Katie, for inviting me to on on this podcast. And yeah, really excited to ah to share the panel with Makram.
00:02:10
Speaker
I lead DoorDash's experimentation platform currently, but I've been on both sides of experimentation, both using experimentation and building experimentation systems.
00:02:23
Speaker
So I built actually exactly if first version of Intuit's experimentation platform across TurboTax, SweetBooks, Mint.com, and then later on,
00:02:33
Speaker
led an effort to open source it. So one one is one of the first experimentation platforms available out there on on them on the market. And then my problem worked on the later versions of of that platform. And then later on, I led Robinhood and currently DoorDash's experimentation platform. But I also have been on the other side, right? So I've used experimentation and companies such as PayPal, Uber, and Amazon. And yeah what what I've learned is yeah no no matter of which size of experimentation you're on, we all benefit. We all should be using experimentation because it's just you know it's it's good for your business, for the growth of your business.
Community of Experimentation Professionals
00:03:19
Speaker
A hundred percent. And you both know each other. That's why this is our first three person Unite Voices episode. So for anybody watching and they're like, oh, I would really like to do that, but I want to bring somebody.
00:03:35
Speaker
We can have as many people as we want on the show. that's That's the beauty of it. So the experimentation community is pretty small, but y'all have worked together in the past, it right? I'm like, I want to make sure I'm understanding that correctly.
00:03:48
Speaker
but We haven't worked. we I don't think we were overlapping, but it's such, like you mentioned, like it's a small community. Like look at it. We are connected on LinkedIn. We chatted a lot of times. He was before me at Intuit managing the early version of experimentation.
00:04:05
Speaker
the people I worked with like Credit Karma, Darwin, experiment, you know even at Intuit there were two versions. There's Darwin for experimentation. So that person is now under Elia's team and Like it's a small world and this is the beauty about the experimentation community.
00:04:24
Speaker
Yeah, exactly. To my point, we we didn't directly work work together, but it seems like we we did big because throughout throughout the years, yeah, we we kind of used each other's work or work with folks who were in in the in the community. but yeah, what what I found is experimentation, and it's it's like the smallest big tech function. It's, you know, the community is, you know, is is is tiny and there are only only so many kind of experts in experimentation out out there.
00:04:57
Speaker
But our work is used, you know, throughou throughout companies, right? It's in an essential part of, you know, big tech companies as as well as, you know. Smaller as more as startups, mid-sized companies. But we all share our learnings about experimentation. And i would say we agree on maybe 80% of things, you know, p-values, variance reduction techniques, experimentation methodologies. And then we argue on on the 10 to 20%. of the corner cases. But you know our our goal is no matter ah where we go, is to build kind of the best platform as well as the culture of experimentation, you know not no matter where we go.
00:05:45
Speaker
Yes. I think that's what I want the audience to take away is I didn't even realize that y'all never overlapped, but it is all in the spirit of experimentation to work together to make things
Challenges with Complex Tech Stacks
00:05:59
Speaker
better. So this question i think is for both of you because you both led such huge high stakes programs. Yeah.
00:06:07
Speaker
The thing that a lot of people run into is inheriting different tech stacks. And i mean, Ilya, I know acquisitions happen all the time. Macrom, you have built programs from the ground up. I would love to hear...
00:06:22
Speaker
a little bit more about the challenge or mistakes that you often see, anything related to what it looks like when you have a really difficult tech stack and what your first step is in what I like to call untangling the spaghetti of multi-platform experimentation practices. Because I know so many people listening probably have a tool for ideation, a tool for build, a tool for analyze, whatever it is. They're hosting everything in Airtable and then putting it all in Google google Sheets, right? Just tell us a little bit more about the more technical side of what it looks like to untangle the spaghetti of ah either inheriting a really complicated tech stack or start maybe you're starting fresh. Let
Integrating Acquired Companies at DoorDash
00:07:05
Speaker
At DoorDash, so we we have a pretty mature platform that we're building for over six to seven years now. But then DoorDash has been acquiring companies around the world you know for our global growth because we've been growing but in both our marketplaces and domains. So going beyond restaurants into grocery and electronics and retail and clothing deliveries. But then for our global growth, we have acquired a company called Volt based in Finland and Dilovaroo based in the UK. And they also have their own experimentation platform. So what I found is that it's it's important and not not to just rip in their place because that destroys the kind of the local trust although of those you companies that are really doing a great job in their own marketplaces. For example, Volt and Deliveroo are really great at localization. you know for for local languages and and custom features that are really specific to those countries and regions. But what we can all agree on to start with is on the methodology as well as metrics that should be used kind of across the board.
00:08:27
Speaker
For example, because we for a delivery space, we're a marketplace across consumers. dashers and couriers and merchants. so And we all agree that for consumers, we want to have the the best quality, reliability, satisfaction of orders that they're getting. On the courier side, we all agree that we need we need to have the best earnings for couriers, as well as metrics such as their utilization, make sure and they're not secure and idle, and finally, fair and resource deliveries, right so everybody gets you know their share of deliveries. Finally, on the merchant side, restaurants or store owners, we are focusing on their order volume, their unit economics of the specific items that are being delivered, and and and finally, kind of their longer term longevity of their businesses, not just focusing on short term, but rather on.
00:09:27
Speaker
But then we allow local you know brands to to actually go initerate and on those common common metrics. But then only and now, after a few years, we're actually consolidating and building a unified platform, but always starting kind of the fundamental foundations first.
00:09:47
Speaker
I love that. I want to make that, you know, part of the subtitle of what this episode ultimately is about because I i do see oftentimes in my personal career, you know, you come into a new program and it's really easy to blame it on the tools.
00:10:04
Speaker
It's very easy to go. That's the problem. But the alignment of the methodology is so important. So Macram, in your experience, are you do you tackle the same problem first when you're inheriting complex tech complex tech stacks?
00:10:16
Speaker
Yeah, exactly. I mean, just like I mentioned before, Darwin. So Intuit acquired Credit Karma. Credit Karma had Darwin as the primarily experimentation platform. And obviously, you know, want to make sure that we are not disrupting Credit Karma business. If this is working, the team is working with that experimentation platform, all is good and dandy. So we wanted to make sure that that is perfect.
00:10:42
Speaker
Now we started to integrate, do more cross-perm work between Credit Karma and TurboTax. You're seeing a lot of connectivity over there and there are experiments that are now running across either starting in Credit Karma and effectively driving new acquisition to TurboTax or vice versa.
00:11:05
Speaker
So now these require now the two tools to handshake between one another. And this is where the discussion started to happen between the Darwin platform team and Intuit experimentation platform team. How are we going to make that happen so that people who do not want to launch two experiments and people are starting to flip-flop between control and treatment between these two variants? So this is where we started that conversation.
00:11:30
Speaker
So you can see like and of each one has its own tech stack. So that's where, you know, the tricky situation start to happen. And obviously there are other reasons that from a company ah combining technologies and stack stacks that are also big effect and factors.
00:11:48
Speaker
You both have such specific experience in this part of the field.
Building Trust in Large Teams
00:11:53
Speaker
I'm really interested in in something that i think a lot of people struggle with, which is how do you share this information and build trust, right? Because you're saying we're aligning on methodology,
00:12:07
Speaker
Like, how do you build trust with such a large team when you have people coming from such different points of view? Usually, in my experience, people feel very passionately about experimentation, which is just incredible, right? What how what do you do when somebody says, oh, I don't agree that's the right metric? Oh, I do actually think it's the tool. I'm curious how you build trust and and culture of experimentation with your teams. Yeah.
00:12:31
Speaker
Yeah, so the way i I like to approach it is through, well, first of all, learnings, right? So it's it's it's not really about the the tools, but about people and processes that that are using this tool. So it's really important to ah first enable experimentation and and show, no matter you know if you have some somebody new joining, to to show, okay, these are the experiments that have run in the past. So so we have the culture that all experiment readouts or experiment signatories, either a win or loss, are sent out across the company. And anybody has who who wants to understand what what happened, they can read it and and learn from it. The second letter is using AI to bring out what's called institutional knowledge, right? We're using AI to now mine all the past experiments readouts, and and you can actually type in a query such as, you know, show me all the dash pass experiments that moved our global order volume by more than 0.2%,
00:13:48
Speaker
And you you will see all those experiments and you can really understand and mine further to to understand the details there. And last but not least, having and what's called leadership and executive buy-in on Ryan experiments and also understanding, kind of replying to experiments. For example, our CEO at DoorDash replies to all all experiments that are are being run and encourages to look at alternative ways of approaching this problem. So this kind of builds trust.
00:14:20
Speaker
So no matter ah kind of what culture of of experimentation you came from, you should be able to kind of understand the past, understand what's happening right now. But of course, you know provide your own input and feedback on what happened. But our goal ah at the end of the day, and no matter what experience you're coming from, is to build the best in class experimentation that really works for DoorDash and our global business.
00:14:46
Speaker
I'm going to also give you a few examples. I entered really the world of experimentation at LinkedIn when like I suddenly, you know, I was a product manager, developer productivity, and then I got the opportunity to manage TRX LinkedIn's experimentation platform. And I worked under Yagzoo, you know, she she worked with Loneca Javi on the book. So that's where, you know, the,
00:15:11
Speaker
Once I became a PM for T-Rex, LinkedIn's experimentation platform, the world of... I started getting bombarded, and trust me, the word bombarded, left and right with feature requests from the different teams, and started to see like the importance of experimentation to the different various teams. So I started realizing that you have AI engineers, data scientists,
00:15:36
Speaker
back-end engineers, why does back-end engineers need, have this need for experiment? Well, because of feature flags and all of the cleanup and all of those.
00:15:47
Speaker
And the same thing, you know, so I had to really develop a prioritization framework because everybody wants their feature yesterday. And we are a small team. How can we like support all these crazy use cases that are coming our way? i actually did a product school talk specifically about prioritization because like it was crazy abnormal the demand that is happening on experimentation teams.
00:16:14
Speaker
Then come, I thought, okay, I moved to into it. Hopefully, you know, less demand. Turns out it's even worse. We have, I'm dealing with B2B use cases for QuickBooks and then B2C use cases for TurboTax.
00:16:28
Speaker
and So not only that, you have, as a, you have, Web authors, the CMS editors who want to, and then you have designers. Then you have developers, front-end developers and back-end developers. So all of these have different needs from the experimentation platform. And as you mentioned, it's easy to blame the tool.
00:16:47
Speaker
And I was managing experimentation platform, personalization platform, and the CMS platform. So it's easy also to start pointing, oh these tools don't talk to each other. Because remember, experimentation is is a piece of the puzzle.
00:17:03
Speaker
For example, if I am running a a pre-authenticated marketing website experiment, well, that touches the CMS platform and then touches the experimentation behind it. So either, and then there's always that, oh o who owns that? Is it the CMS or is it the experimentation?
00:17:23
Speaker
So eventually to tackle all of these, we i even arranged Go-To-Market Tech Summit, an annual summit event. where we invited all the key stakeholders from Australia, from London, and became that get together on a week on a week where everyone will come to Mountain View, and they will be talking about their issues.
00:17:49
Speaker
And we even like planned it ahead of time. So they, one week before, come prepared. All the positive things you have to do to say about us so that we're not like, oh, I hate this week, I hate this week,
00:18:01
Speaker
You know, something is working. Tell us what's working. again And then come back with the things and you pick three because you have 20. Pick three and let's see how we can tackle them.
00:18:14
Speaker
So that really relieved things. So sometimes, you know, these human interactions are very important. I think that's an interesting juxtaposition in your answers, right? It's Macram saying bringing people physically to the table. And Ilya, something you brought up is leveraging AI to democratize insights. And that's something I talk about a lot.
00:18:34
Speaker
But I think both pieces of the puzzle create trust. And I'm curious, you know, I realize this is a little bit tangent tangential to our conversation, but I want to make sure I'm asking this.
Leveraging AI in Experimentation
00:18:46
Speaker
Macrum and I talk about AI all the time, right? Macrum is a strategic advisor to Chameleon, for everybody who didn't know. Macrum and I were are working on a PBX 2.0 series this morning, and we've been working on it for weeks now. So I know how Macrum feels about where AI goes into the workflow, but I may have you repeat yourself on this podcast, Macrum, because I think what you have to say is really important. So Ilya, I'm curious, where are you leveraging ai Where is it providing value? and where, do humans fall in the loop still?
00:19:20
Speaker
Yeah, yeah. Yeah, we are. Well, I would say we are starting right now on the ai journey. But but yeah, we ah we've launched what we call is a MCP server for experimentation. And it's, you know, it's it's it's been going viral. and p People have been, you know,
00:19:38
Speaker
building their own skills for the various stages of experimentation life cycle so where we're trying to you know kind of vet and contain it to make sure that people are not carry cutting in corners and they're still using the established standards of experimentation but but yeah we have a is a skill for example to create an experiment or a feature flag We also just recently built out a experiment summary and readout skill, which kind of summarizes all the experiment analysis results. So it's really and to understand what what happened to experiment.
00:20:10
Speaker
But also we've been building and kind of a background ai agent that ah looks and scans for your repositories and looks for stale feature flags or experiments that you know that you should actually remove because it's know it's called that code.
00:20:29
Speaker
But what I'm realizing is we need to build a lot more skills throughout the experimentation lifecycle. For example, at the design stage where you know you no longer need to wait for an experimenter to create a Google document with with their definition, they can start and actually design their experiment right right there.
00:20:50
Speaker
with with that AI agent, with our hypothesis and what metrics should be used and things like that. And also helping with those background agents to debug those experiments that have gone wrong, for example, with sample ratio mismatch issues or you know anything else that needs space specific, maybe restart of experiments. So yeah, we've been introducing AI throughout experimentation lifecycle, but at the same time, we still want to have some you know UI that's available though where they can look at the experiment analysis results. But you know more and more people are just you know using their agent command line to to really interact with their experiment.
00:21:34
Speaker
Before a macrum, I let you go into the spiel that I know you're so good at. is Isn't this funny? It sounds like Ilya was in our call this morning in terms of looking at stale feature flags. You and I both went that, that stale feature flag, because that's something we were talking about this morning because our Chameleons PBX 2.0 is launching. And by the time this show is out, actually, it will have been launched. So spoiler to everybody. This was prerecorded, if you didn't realize. And, um, but it's, it's PBX 2.0. It, a huge part of it is making sure you're taking it. You're taking advantage of your test winnings in a reasonable time. So you don't have stale feature flags. So you don't have stale learnings, whatever it be. I think it's hilarious. We're going to send you what Macram and I recorded this morning, and you will hear a lot of the terminology you just used because we were, the whole thing was about AI, but Macram, I, I want you to tell us a little bit more about what you think, where you think AI provides the most value in the experimentation workflow today?
Phases of AI Impact on Experimentation
00:22:31
Speaker
I see three phases of AI hitting us right now. Phase one, just like Elia described, it is like streamlining the end-to-end workflow from ID8 to shipping.
00:22:44
Speaker
That should be all, know the less dependencies. Right now, like if we talk about six months ago, we had two bottlenecks or one year ago.
00:22:55
Speaker
One bottleneck is developers because every experiment needs developers to build, set up of that experiment. There's a lot of setup time that is doing nearre required. And then there is also launch time after we we have a winning experiment to get it out of the door. That's why we have all of these forever running experiments because nobody has a bandwidth to turn it off and make it right and all of these feature flags and all of those.
00:23:23
Speaker
So there is a bottleneck of engineering. There's another bottleneck on analysts or them to really do the readouts. And the analysts not only on the readout side, but also the experiment design side, because in many cases, if an engineer doesn't set it up properly, then garbage in, garbage out, that experiment completely is not the proper.
00:23:45
Speaker
So how can we remove these bottlenecks and have with AI and much more and democratize the experimentation so that not only experts in experimentation are only running experiments. Ideally, we want the tool to be an intelligent tool that anybody, a designer, a marketer, a PM who's not technical is able to go and run with it.
00:24:10
Speaker
A designer coming from a Figma design is able to just push their Figma design into a prompt and get it up and running because the whole notion of experimentation, let's test it out.
00:24:25
Speaker
you know I get into so many meetings, oh, this is gonna take a lot of energy. like I'm not asking you to go to production yet. I need to just test it out to see if it's working and then we'll talk about that piece.
00:24:38
Speaker
And this is where the beauty about you know how can we remove the setup time, quickly launch it and then get the analysis. And then if it's winning, how can we quickly No dead time because to get it out or to ship it out.
00:24:53
Speaker
So that's phase one in my opinion. Phase two, we're going to hit the um traffic bottleneck now. Our websites, unless, you know, of course, DoorDash, you have on the homepage of DoorDash, you have the luxury with traffic.
00:25:08
Speaker
But we need to be able to do testing everywhere, not only the homepage of DoorDash. The teams who are owning the different components of the, they won't have a lot of the traffic. So how can we enable them to do other flavors of A-B testing? Maybe not online controlled A-B testing, not the causal stuff.
00:25:27
Speaker
So one thing that I was doing at Intuit is bringing the toolkit of experimentation so that they can do early testing. Early testing is better than not testing. At least they have a better data-driven decision on that.
00:25:41
Speaker
And then the third phase is synthetic audiences. There's a lot of debate right now about whether synthetic audiences are going to are scientifically proven, but like I mentioned, and if you are caught on time and you only have a small period, I mean, just let's take an example, TurboTax.
00:25:59
Speaker
TurboTax, all year long, we finished taxes now. Do you know how much traffic is going hit TurboTax.com? It's going to be not too much. When is the TurboTax.com going to hit I imagine from January till April, right?
00:26:14
Speaker
So, but that's showtime. And we do not want to be doing a lot of bad tests during that period. So how can we leverage synthetic audience during this slow season and do with synthetic audiences something good so that we are running those tests as well.
00:26:38
Speaker
So you see, that is error number three in my opinion. Love this. Showtime. I can imagine that really is showtime. We could spend an entire other 30 minutes talking about ai and the future of experimentation. um I will leave us with a quote I saw on LinkedIn today from Johnny Longden that said,
00:27:00
Speaker
A-B b testing will not survive ai but experimentation will know
Advice on Simplifying Experimentation Programs
00:27:06
Speaker
the difference. And I was like, oh, I love that. Like, probably true. So interesting. We could talk about this all day. i just know there's going to be people listening, looking for your leadership on everything. But at least for now, they've heard about how to navigate really complex tech stacks. They've heard about how to build trust. where you're some of the leaders in the space are seeing AI go or where you're using it now.
00:27:31
Speaker
I realize we're coming up on time. So I want to, I always end every episode with the same question. I want to ask you this. If somebody is leaving this episode right now and they are so overwhelmed with a team that maybe let's say they're just not understanding where they can create simplification, because I think that's what you do best as leaders, right? Both of you have shown that here is you're able to take a very complex issue and make it as simple as possible to create more impact and And whether it's simple just to the receiver, let that be it. But if you could just give any advice to anybody who wants to create simplicity in their program, whether it be tech stack, whether it be trust metrics, whatever it is, what would you recommend they do tomorrow?
00:28:26
Speaker
I can start, I would say, you know, kind of what I learned throughout my career is earlier on, I focused on skill
Focus on Decision Quality over Quantity
00:28:37
Speaker
ability. So I was talking, and oh, how can we run more and more experiments?
00:28:42
Speaker
and kind of And quantity key is is the goal because the more experiments you're on, probably the but the better things will be. But what I learned is it's it's not it's not really about the the quantity, but more about quality.
00:28:58
Speaker
right You can be running a lot of different experiments, but if nobody is really looking at the results and it's not really changing your initial intention of what you're going to do, and you know you can just look at experiment results and launch your feature anyways, even if it's the great in your guardrail metrics and things like that, you know so so so what? So instead of pursuing this magic number of 10,000 experiments, 20,000 experiments, experiments, focus more on decision quality.
00:29:30
Speaker
So how many original calls that you are going to make it actually made, right? So looking looking at at experimentation as a decision engine, this is why, for example, my group at DoorDash experimentation group, we actually call ourselves our decision systems. So we we are in the business of not just trying experiments or not enabling feature flags, but in driving decision making. So, yeah, full focus on decision making, not on ah just pure reporting of of experiments just for the reporting sake.
00:30:10
Speaker
I love it. I love it. You know, you talk very similar language to me. The more I listen to you, Elia, the more like we are very, lot of things in common for us. I love it.
00:30:20
Speaker
These vanity metrics that we typically chase, I've been in many of those meetings. Like, okay, what are we trying to achieve here? What are we doing here? I love it.
Embracing AI for Improved Experimentation
00:30:31
Speaker
ah From my side, embrace AI. We live in a different world now. And then if you think about the Venn diagram where you know PM is only doing PM work, designer is only design work, that one is going to coming together.
00:30:50
Speaker
And at IDME especially right now, a engineers are writing product PRDs, they're not waiting for a PM. If that's space that they can do it, let them do it. We are democratizing, everybody can write content Like everybody is now able to launch an experiment.
00:31:11
Speaker
It is those kind of with the right checks and balances. I can check in code. We have the rigor on the PR request. Like it's not like ah only an engineer can check in code because of the pull request and all that reviews. Man, everybody's using these AI tools now and we can check in code.
00:31:30
Speaker
The auto PR with the right rules and we have genetic architect and agentic PM that we're putting them together. So we are streamlining those. So those old habits are, you know, those are the difficult things. If you can get rid of those habits and embrace the new world, that's my advice.
00:31:52
Speaker
Overall, wonderful advice from both of you. We have be a decision system and embrace AI. I think that you can use AI to streamline your decision systems. Also, for anybody watching on YouTube, you're getting a really good view of my And my cat muted me.
00:32:12
Speaker
Yeah, my cat is truly laying directly on my keyboard. It's fine. um I'll recap what I said, which is just great feedback for everybody. i think they have a lot to learn from both of you. i wish we could have covered all the topics that we wanted to cover in 30 minutes, but it's just impossible. There's so much that you both could say. And I would encourage anybody following along to maybe find you on LinkedIn. If you have a question, ask them. But really good leaders to learn from and leveraging AI and being a decision system, really important pieces of the puzzle. um I would say be cautious of how you're leveraging AI to make decisions, right? That's my piece, if I'm going to add my flavor to it, is
00:32:56
Speaker
AI is an amplifier, not a strategy. and you are the strategy that employs AI. So I think like bringing those two together, I'll bridge it there.
00:33:07
Speaker
Really perfect way to pick up tomorrow and start. But thank you both for your time today. i really appreciate you. And yeah, I hope if anybody's watching on YouTube, they're getting a really good view of my cat. But in the meantime, we'll see you next time. Thanks for tuning in to Unite Voices.
00:33:26
Speaker
Thank you. Thank you so much. And as as they as they say what it was called, human human is always at the helm. right AI system, but the person, humans should be making the final final calls.
00:33:38
Speaker
Oh, well, there's the title of the episode. Decision systems, humans at the helm. Love it. It's so great. Thank you both.