Transcript
Speaker: Yeah, exactly. So everything sort of flowed from, okay, what can we how do we get to somewhat work in sentences in one day? Like I tried it yesterday. I trained another small model just to fund on the latest version.
Speaker: And I only gave it a day and it's saying really cute things, but it's like, I asked him, what is the sun? And it's like, oh, the sun is a moon at the bottom of the earth, which is a real sentence. That's actually kind of like, is that deep or is that just dumb?
Speaker: It's just dumb. That's something a thing a two-year-old would say. It's like, that's adorable. Well, like the cool thing is that it's like even understanding that I'm asking about the sun and it's giving me like even like eat the fact that I understand that I'm asking it a question I should answer.
Speaker: like Hey friends, I'm Scott Hanselman and it's another episode of Hansel Minutes. Today I have the pleasure of chatting with Felix Rieseberg. He has a long history of building delight for developer tools, tools that you have used. He's worked at Slack, he worked on Electron, he worked at Microsoft, and his recent side projects have become very popular because they make AI feel approachable rather than just cha chasing benchmark scores. And his latest language model builder is in that tradition. It's basically an educational app but lets you build and understand language models themselves.
Speaker: How are you, sir? I'm doing well. Thank you, Scott. Thanks for having me on. Always wanted to come to the show. I appreciate that. I should have had you on many, many years ago um you know because we worked together for a time. i always love the stuff that you do. Sometimes you'll just do weird, delightful stuff. you're just like, oh, look, here's Clippy. Here's Windows 95 running in a browser. like Why? Because it's delightful for no other reason.
Speaker: But, you know, you work, you've been and working in open source forever. You worked at Microsoft as an open source engineer. You tried to make people's lives better by putting things on GitHub. So you always come at things from a humanistic perspective, which I appreciate. Now you do in your day job work at Anthropic, but this, this podcast is not about Anthropic. This is about you and the cool stuff that you're building, but we do want to Shout out that you do lead engineering for Cloud AI for co-working for code desktop. So maybe we'll have another show another day.
Speaker: But this is about something else. There's a guy but named Ben Eater who makes a 6502 create your own processor from scratch. Yeah.
Speaker: Okay? And he's like, hey, let's go and get these, like, we're going to build it from scratch. Like, we're going two rocks and we're going spark them together. We're going to get 7400 series parts. We're going to solder it. And at the end of this crazy experiment, you will have a computer that can say, hello, world.
Speaker: Yeah. that's I feel like you did that. I feel like you did that with Language Model Builder. That's, like, the nicest thing anyone has said to me, like, this week. I think think the...
Speaker: I think that's sort of the goal, right? Like I looked at Language Model Builder because I've been i've been in ai for like a little bit now. And you mentioned that we were Microsoft together. I don't know if you remember Project Oxford, like from that. yeah.
Speaker: Yeah, there was like an Ngram model in there. And back then I was super junior at Microsoft, but I made like the NPM modules for that Ngram model. And you could give it like worldwide. It would think for five minutes and it would come back with a worldwide web. Maybe that's like a good completion of like what you just gave me.
Speaker: Yeah. And um obviously, like AI has moved pretty dramatically, but I find myself like explaining the thing to people a lot. And I think one thing that is cool about AI is that the for me, the perceived delta between how I don't want to say simple, but the core ideas, I think, actually quite simple and very understandable to how magical they are is like so big.
Speaker: that I often find myself sitting sitting people down and being like, right, let's take a napkin. I'm going to show you how a language model works. And I think for me, at least, that that feels so much easier than explaining to someone how like a modern chip works.
Speaker: I have no idea how modern chip works. To me, it's like all black magic. right like we're like Sometimes I look at the nanometer diagrams and I'm like, this is this seems hard to explain. It's easy to explain.
Speaker: There was a cartoon from a guy called, ah it was called The Far Side. It was like a cartoon that they would have on the Sunday newspaper. And there's a two scientists sitting in front of a whiteboard and they've got a lunt bunch of complicated calculations and math symbols and then a little sparkle in the middle. It looks a lot like the clawed sparkle.
Speaker: And then it says a miracle. And then on the other, on the right of it, there's like a bunch of other complicated, you know, math. And the other scientist says, I think you need to be a little more specific here in step two.
Speaker: I feel like that's what's going on with language models right now is that like we can understand the number 10. We can then times that by 10 and then times that by 10. But then you add a couple of orders of magnitude and then we just go and then a miracle happens.
Speaker: Yeah. And now the the magic word guesser feels like it's a person. Yeah. Yeah, exactly. And I think it's like so cool, especially i think you mentioned my day job. Like my day job, I get the privilege of like,
Speaker: you know, seeing baby clot in the oven and like just seeing it evolve and grow. And you get the, the, the joy of watching something just like put together random words.
Speaker: And like, you get these like complete word salads and you get like this complete noise of characters back and like watching that in sampling go from just, know, you know, face planting onto your keyboard to like, oh, I recognize some of those sounds to like, oh, now a sentence is happening. It's pretty cool. It's like very fun. It's like very fun to like talk to something that you made.
Speaker: Even if it's just a model, it's just like a deeply enjoyable experience. But is there, isn't there a leap? Isn't there a moment there where it's just like that went from cool, cool next token guesser to this feels smarter than a parrot?
Speaker: That's the part I don't understand. How do we get from word guesser to parrot to people are talking to this thing like it's their therapist, which is concerning. Yeah, i think I think there you go into these like deeply... God, I gotta be i got to be fair careful not to be too annoying in this. I came to computing fairly late. I studied poetry. I at some point discovered that the poetry factory was not hiring and I needed to make money some other way.
Speaker: And then I guess now i' building software. But I think um language models give us such an interesting opportunity to talk about all of these big concepts, right? Like we already have these big conversations about moral patienthood, like at what point do we even consider seriously whether or not moral patienthood. But it also gives us these really interesting ideas about like, you know, like what does it mean for something to be intelligent? Like clearly we know how it works.
Speaker: If you get meaningful therapeutic insight, what what do you do with that? right like And I think the questions there are probably much bigger than like, how does the language model work? um And they're more about like, are we all just parrots? like Where does it what is it break down, right? Yeah.
Speaker: Yeah, I try to explain context windows to people. And the sample the the example I was giving my talks is, it's a beautiful day, let's go to the. And then the model always picks park or beach.
Speaker: And then you give it a little context. You tell that you're in London. It'll say, let's go to the London Eye. You tell it you're in Berlin. Let's go to here. You tell it you're in Portland. Let's go there. And the more context, the better the word guessing gets. And then you say, well, you're Scott Hanselman.
Speaker: And you've been married for 25 years. It's a Saturday afternoon. And the answer is that we'll go to the bagel shop because that's where my wife and I go for the last 25 years. Oh, beautiful. That's our Saturday tradition.
Speaker: So then the question is, am I just a big corpus of 25 years of her putting up with me? And she knows to guess that because she's the word guesser for my personality. And am I just making stuff up?
Speaker: Yeah. Or is there is there is there free will? And then it's a whole thing. Yeah, yeah, like you very quickly, you very quickly Or I at least, this is like a very personal experience, but I think it's like part of the experience that had me both this language model builder thing. But like when I stare at the model, like the big models, I sometimes get the same feeling that I get when I like stare at the campfire or like a forest or like the stars. Yeah. Right? Like this notion of, i think people have different experiences, but like the experience I have when I look at, when I go to a planetary, I love the planetary. I just like sit down and I feel so small and insignificant and it really makes me
Speaker: consider the uniqueness of the human experience and like, honestly, the big language models, when I look at when i look at this like statistical, mo like just math, right? Like when we look at a language model,
Speaker: What I mean here is I'm like less excited maybe about a particular application or like a particularly trained model. But like if I look at a model as the core idea of we have math and I'm talking to the math, mean I sort of feel like the universe is like staring back at me. And that is a very cool experience because it makes me reconsider myself and like, how do I put work together? You know? Right.
Speaker: Well, Language Model Builder is at languagemodelbuilder.com. You can just download it. Now, right now it only works on a Mac and it expects Apple Silicon. So be aware. And you you start with a kind of an interactive book.
Speaker: You know, it's it's like a it's more than a PDF. Everything can be clicked on. Everything can be changed. every It's kind of like the best of the New York Times when you go to one of those New York Times infographics and you're like, oh, man, that's not a GIF.
Speaker: I can click on that. Holy crap. That's awesome. i'm going to I'm going to print some of those quotes and just hang them in my office. that Oh, you can take it all. My show is your show, my friend. So if you want me to blurb blurb your product, you can blurb me as much as you want.
Speaker: Now, I want to call out though, when you get to the tokenizer part and you say, I'm going to write my own tokenizer, byte level tokenizer, you might just have a vocabulary of 10,000 words. So that might be like a baby, right?
Speaker: Yeah. And then you can make like a GPT-2. And I think it's important for people to remember that like GPT-5, there were four others before that. and they did stuff, right? So you're basically putting people just before the bendy part of the hockey stick of the graph, right? You're saying right before everything exploded, this is the size of a model we can make where the human mind can kind of still get it.
Speaker: Which is why I use that Ben Eater thing as an example, because a human being can completely and 100% understand a 6502 microprocessor. They cannot understand a Pentium and there's no way they can understand like Apple Silicon and and hold it in their context window.
Speaker: Yeah. How did you decide the right size for language model builder to be so that you could understand it all? Because I think you succeeded in doing that. um I think it was probably like less strategic than it might seem like now that we look at it. But it was I was basically working backwards from, okay, my friends, my family, like what kind of attention span can I assume?
Speaker: five What can we train within that time? And the tokenizer is a good example because I actually, i eventually settled on the like 10K tokenizer. I'm the one that is like very small, but mostly because running it on my my personal computer, so like an M4 Max, which I think is like a pretty good computer, all thing all things considered.
Speaker: And with even with the later later tokenizers, you spend so much time just tokenizing the corpus that you just like stare you're just like you know you're just staring at the pre-training, not even pre-training, just spinning its wheels on the tokenizer.
Speaker: It's just like mashing up a food for a baby. You're just sitting there mashing up potatoes. Yeah, exactly. So everything sort of flowed from, okay, what can we how do we get to somewhat work in sentences in one day?
Speaker: Like I tried yesterday. I trained another small model just to fund on the latest version and I only gave it a day and it's saying really cute things, but it's like, I asked it, what is the sun? And it's like, oh, the sun is a moon at the bottom of the earth.
Speaker: Which is a real sentence. That's actually kind of like, is that deep or is that just dumb? It's just dumb. That's something a thing a two-year-old would say. It's like, that's adorable.
Speaker: Well, like the cool thing is that it's like even understanding that I'm asking what the son and it's giving me like even like e the fact that i a understand that I'm asking it a question I should answer. i' like That alone is like. Or that it's a round thing in the sky somewhere. Like it's got,
Speaker: it yeah the the the the vector space is at least has some amount of proximity where it's like, okay, moon and sun are near each other in the space. Yeah. And then also the other thing that happened is like, I sort of wrote like three different versions of this app. The first version was just the textbook. And what I kind of wanted was, you mentioned these like playgrounds, but there's like the little playgrounds of like, okay, let's give me your corpus. We'll make a tokenizer together. And the initial idea I had was, oh, at the end of that textbook, you have a model.
Speaker: But then very quickly, I found myself like wanting just a little bit of innovation that we found in the industry. So at some point, I made the call to like separate out the playgrounds and say, okay, we're going to learn the foundations and then we're going to build a model.
Speaker: um But I didn't want the gap to be like too big, right just so that what you're learning is still applicable. I guess you're right. Tokenizer is a great example. like The modern tokenizers, um they're just little there's just a little too for much magical science going on for a human to quickly honest understand.
Speaker: I have an analogy that I use that everyone is sick of hearing, but like learning to drive stick shift, manual shift on a car changes your relationship with the vehicle.
Speaker: Yeah. Language model builder would by its nature change your relationship with a chatbot or with a large language model. Because now suddenly you're like, before we get to talk to the bot, we're going to take apart this car and we're going to talk about the internal combustion engine and the history of that.
Speaker: you don't go too deep into the history. You don't talk about like 40 years ago, Eliza, you know, how did you balance? There's the pure math there. You just kind of like assume transformers are a thing, but you don't really explain where it came from the last 50 years. Did you think about historical context versus mathematical context versus just explaining the moment that we're in today?
Speaker: Yeah. Like just a little bit. I think I wanted to keep it very tight because um On the website, I wrote the sentence that I think is very clever. I wrote no artificial additives, by which I basically mean, obviously you can sit down and you can say, Claude, write me a five volume book about the history of AI and and would do it for you.
Speaker: So I wanted to to stay like very true to the things that I, like if if I sit down with a smart friend, right, like maybe like in medicine or like law or something, I sit down with a smart friend and they're like, hey, can you explain to me how language models work and we have one dinner?
Speaker: Like what would I put in that one dinner? So I gave myself this like somewhat artificial limit of 90 minutes. And within those 30 minutes, you can touch on just enough history to like understand the current context. right like Why do we give models a temperature? Where does that come from? What is a hot model? What is a cold model?
Speaker: So we go into that. We also briefly go into like what happened in 2017, like what is going on with like attention is all you need. Like it's famous enough as a paper that people have heard about it, but i'm I'm not spending too much time on like um statistical models of language.
Speaker: um I think it's a super interesting history. I really enjoy like,
Speaker: Let me figure out how to put this. I really enjoy, I never studied linguistics, but I hung out with a lot of like linguists and like computational linguists were like scraping at the surface of AI for like the last 40 years. And like, we didn't even know, right? Like I studied with those people. they kind things work cuts it just like Well, it's it's the elephant problem, right? You're running around and you're like feeling it. And the one guy is like, this elephant is, this this is the part of the blind people touching an elephant. This is the this elephant is like a tree.
Speaker: No, no, an elephant's like a snake, but no one sees the big picture. Yeah, yeah. There's like so much cool stuff there. And I think i think some of the, I'm sure we're gonna get like amazing books about the history of AI at some point.
Speaker: But you don't think it's important to understand? like that You can understand language models this way. So how long is this? Is this a five-day course? Is this a 90-minute course? It can be as long as you want it to be, but would I give this to my 20-year-old and would that be his introduction to StickShift and change his relationship with the chatbot? Or does this require someone to really have a background as a software engineer?
Speaker: now No, no background as a software engineer. There's these little drawers where you can look at the code if you really want to. code isn't even the real code. What I'm using is MLX, the Abno framework. That's why it's currently only running on on Macs because I was just lazy. and I'll show you some Python if you're interested in like what would the Python look like.
Speaker: Yeah, because I wanted to know when you're going to get this working on Windows. Yeah. um I mean, honestly, at this point, I could probably just like ah have an agent do it, right? It shouldn't be hard. You can just do it in Python. Well, but then you got to get it to work on ARM and X64 and WCell or not WCell. Like you have to like the combinatorics get complicated once you decide to make it go outside of the homogeneity of this Mac.
Speaker: I didn't want to, I didn't do it a priori for like a few reasons. One is like, I think I didn't want to build yet another Electron app. i wanted to like actually build a Swift app for phones. Yeah.
Speaker: yeah Yeah. That's a good point. You'd have to do it when UI. Yeah. Yeah. But the other thing is like, you you never know if like, am I really only building this for my five friends or does anyone care? People do care, which I think is very nice. And I think that increases the likelihood of it going up.
Speaker: But to answer your question, if people have like one afternoon, you will have a model that like will be very stupid, extremely stupid, but you can make it through the whole thing in like an afternoon. That's really powerful. Like this is significant. And I feel like people need to understand that. that And the at the root of it is this concept that it changes your relationship with the vehicle.
Speaker: Yeah, exactly. Opening task manager, opening task manager on op on a PC, opening a DAW. Like I've seen DOS prompt. You go, oh, open a DOS prompt. Like, what is this? I've never seen this before. I've been a Windows user for 20 years. I've never seen. Well, that's, you or hitting F12 on like Chrome.
Speaker: Yeah. Those moments are a moment where like the muggles change their relationship with the thing. and theyre That's like opening the trunk, opening the boot rather on the car and like look at the engine. Like what? There's a thing under there?
Speaker: There's two things in there that I think people should try out because I think they're very fun to me. The first one is I built this like visualizer of what a model architecture looks like. Like what are the pieces in a model? The model blueprint?
Speaker: Yeah, exactly. They're like transformer blocks and like how does how does like the math flow through there. But the other thing i really enjoy is at the very end, but you can also do it in sampling, but it's the x-ray view. So I think a lot of people know now that a language model is fundamentally a statistical model and tries to predict what the next token is. The next token being like you know like a word fragment and um it doesn't pick the most likely one. It's sort of just like a soft min-max and But it it's sort of like tries to calculate that across like all the text you give it and then continues. And what I've built in is this like x-ray view where you can see that every single position, what are the paths not not taken?
Speaker: do you get I call those parallel universes. All the log probes. And it's like, yeah, and there's another parallel universe where it said a totally different thing. Yeah, maybe Scott didn't go to a bagel shop, right? Like he went, like what are the other places he could have gone to?
Speaker: Actually... Is there screen sharing here? like There's no screen sharing. This is an audio show, I'm afraid, my friend. Fair enough. But people need to try it out. um You can do one thing maybe that is very quick.
Speaker: Instead of pre-training your ya model, you can also import an existing base model. like You can like go and import one of the smaller models that just runs fine on your computer, like Quen or something. And you can import that model, go to sampling, or go to the chat interface, but sampling is probably better.
Speaker: And you can say... It's a beautiful day and Scott went too. and Then just see all the possibilities that the model comes up with and like what the properties are. I have a sample app that I use in my talks that draws that as a heat map and then it colors the word based on the log prob and red words are rare words and yellow words are medium rare. So if it'll say, it's a beautiful day, let's go to the, it'll say beach, you'll hover over the B and you'll find out that B and each have been broken into two tokens. You'll click on the B and then it'll pop up all the other log probes for that token. And it'll say beach or B, the letter B is like 75%. And then it'll have like mountains, be like 0.1 or park was 20% chance. And then it's like, yep, at this moment, the universe split in half.
Speaker: Yeah. One fifth of us went to the park and the other rest went to the mountains and this guy went to the beach and that's happening every day, all day, constantly. Yeah. Yeah. And I think that is really cool.
Speaker: It is really cool. Yeah. And to your point about like changing the relationship, right? Like once you, once you get a better understanding of, like the underlying concepts of probabilities, like it makes it easier for you to understand like how models come up with the text that they come up with.
Speaker: But it also helps you understand like a lot of the things that we find most fascinating today about like both moral patienthood and like intelligence and like how is intelligence formed, like at what point do we consider this to be like a useful therapist. like A lot of the things that we're looking at are how do the models i evolve their internal weights and those like transformer blocks, what is happening inside those transformer blocks. And it's much easier for you to get like, for you to join the fun of like all of this, like,
Speaker: you know, armchair philosophy about the universe and like the human experience. If you have a rough idea of when I say transformer block and weights, what does what does that actually mean? um And when researchers talk about not building a model, but growing a model, because it's like much closer to like growing an orchid or something than like writing a tool.
Speaker: If you understand why they say that, like there's so much discourse about AI that is really fun to like read and participate in. Oh, yeah. Well, we should pull a little bit on the thread that you just opened there where you just said moral patienthood, which I want to guess a lot of the people on the call aren't familiar with. And it is grounded in the assumption that someone has a certain amount of cognitive capacity, some certain amount of intelligence. You have to have autonomy and self-awareness.
Speaker: And you have to care about something. The problem is that a large language model doesn't have self-awareness. It doesn't have caring, but it does have value because if the values are the weights and it does have autonomy, or at least it can be given autonomy. So then the question is, you know, is it...
Speaker: does it have moral patienthood but or does it just have the simulacrum of that? Yeah. Yeah. It's like a super fascinating field of study and I'm not qualified even remotely to like, yeah, if you're any useful answers, but I, I really enjoy sitting on the sidelines and like listening to established philosophers and like,
Speaker: um People have been studying like animal ethics for a long time. I really enjoy reading about where they come from. And those people tend to like really, they're quite fascinated about the the individual transformer blocks and the weights and like how the model organizes its own weights and its own blocks and like what kind of function it assigns to blocks, which is never programmed, but sort of like evolves over time.
Speaker: And this evolution of a time as I go through training is' like really fascinating for us because it gives us interesting insights into into um just how intelligence operates, this artificial intelligence. Another useful thing about Language Model Builder, and thinking about that in in the through the lens of some of the language that you just used about how how the weights are applied, how the model thinks, how the model changes, these are all kind of, you're dancing around the anthropomorphizing of these things. But people need to understand, and you can see it when you run Language Model Builder, that like,
Speaker: If you're running a training, you can watch it. You can see it happen live. But once the model is baked and you pull it out of the oven, it doesn't continue to learn. Yeah. Unless you decide to refine it and put it back in the oven and bake it some more.
Speaker: You know what I mean? And I think that regular people and like non-technical parent don't realize that, that it's a bunch of stateless calls, you know, stateless HTTP calls. There's no state here. There's just context given back to you, rehydrated, and then it keeps talking. So these like ongoing persistent conversations that we have are an illusion in themselves of these stateful, the state is is a illusion the state itself of the conversation is an illusion just made by context.
Speaker: Yeah, yeah, yeah, and that's right. Yeah, exactly. It's like the difference between like one thing is like the actual weights, right? And then the other thing is just like additional information or tool costs that we give the model. But the thing that we actually then train is like using other the tools, not not like what to necessarily do with the results.
Speaker: um And obviously, I think at scale, some of the stuff like breaks apart a little bit because a lot of the things that we may have done in like post-training are increasingly moving into pre-training. Mm-hmm. like Especially tool calling, right? like I think there was ah there was a point in time where you could really tell whether or not a model was trained on effectively calling tools or not. It was a brief point in time, especially in coding, where some tools were just incapable of using Bash, even though they could give you like a whole Bash manual.
Speaker: Right. And then in itself, it was little interesting, but... I think that the tool calling stuff is super interesting because when I teach that, I talk about how you are having a chat between you and the model. So there's two entities in the chat.
Speaker: And when you add tool calling, there is a third person in the group chat. Yeah. And then you leave. And now it's the chat bot talking to the tools and the chat bot tells the tool. And then the tool isn't a chat bot, but you can pretend that it is because you give it input and it gives you output that we don't know what we're going to get.
Speaker: And that keeps the conversation going while you're basically on hold as the third person in the group chat. And I don't know if you do this with your people, but like one thing I've always enjoyed doing with with people asking, me okay, how do I build effective tools with AI is to think about like model experience.
Speaker: yeah I'm like, okay, imagine you're the model, right? And like Scott asked you, where should I go today? Okay, you're the model. Like what what kind of what kind of tools do you want, right? Like, and you might be like, okay, I probably want to know like,
Speaker: Is there anything that tells me like Scott's preferences or his location? Right. Do I have memory? Do I know where he is? Is this the first time I've ever met this guy before? That's why the Open Claw stuff is so clever. This simple idea of a soul.md is the every single time you say, hey, are you there?
Speaker: It has to wake up from a dead sleep. stumble out of bed go, what? Who am I? It's like that movie Memento where the guy tattoos all the context all over himself and has to wake up every morning and look in the mirror to figure out who am I and what am I doing here?
Speaker: So every time he learns something, he tattoos himself. Yeah, a little bit. I do think... I think the this stuff all like changes so quickly, but I do wonder... Obviously, most models are actually being trained quite a bit of information about who they are what they're supposed to do. Right. um And I think the open claw thing works pretty well because like most models are just like pre-trained like Obviously, we pre-train on like all the world knowledge, but then we do like a mountain of fine-tuning on...
Speaker: you're supposed to be helpful. right like and give If given this, you should do it. Yeah, that's a great point that you bring that up. Goal-seeking. I keep talking about people that it's not just ah an and so it's not just an intern with unlimited energy.
Speaker: It is a goal-seeking and it has been trained to be helpful. I know that And you can see it and when you go in through your supervised fine tuning, your tuning, your direct preference optimization. and when you go through that in language model builder, the whole point is to make a helpful and useful and kind and goal seeking model. Otherwise it would not be useful. It would just be a sassy bot that would just be mean to you. I did add like the sassy mean bot thing. If you have like a...
Speaker: Awesome. I got to check that point out. Like a thing a thing I added like a little later. um This makes for terrible fine tuning data. Please like I've pre... Don't don't do that. Okay. Like prefix.
Speaker: Prefix. This is very fun to do. And the results are going to be terrible. Okay. Just like putting that out there. But you can import... i added a thing that lets you import an entire group chat as fine tuning data.
Speaker: Hmm. Uh-oh. Then it's going to talk the way that you talk with your friends. That's right. um So it basically like teaches the model, like hey hear are the people, and it's going to continue like the whole continue the script thing, but for like a long group chat.
Speaker: um But it's obviously quite different from the other fine-tuning stuff. right like The other fine-tuning stuff, um and ah obviously, Scott, you've done this many times, but like for the benefit of people who listen to us, um once you have a model that knows about the world, you then have to teach it that if it's given a question, that it shouldn't like give us more questions, but should actually give us an answer.
Speaker: We do that in fine-tuning. And for fine tuning, I have all these different I have all these different examples. There's like one that fine tunes on math problems. There's another one that fine tunes on like being able to rewrite given text is like different text.
Speaker: But I have this one that is like admittedly terrible where you can import a group chat and it will then be like a mimic of that group chat and continue. I made one that talks like ben ah Benedict Cumberbatch and the guy who plays nice Sherlock Holmes.
Speaker: And he's like, well so irritated. Why are you asking these dumb questions? If only you were smart like me, Sherlock Holmes. That's that's beautiful. I mean, I also put in Caparthi's Tiny Shakespeare.
Speaker: um Awesome. It was like one of Caparthi's first demos was whatever language model that writes like Shakespeare. Yeah. Which is very quick to train. But my my point is that like this difference between like waking up and trying to figure out who you are, i think there's already...
Speaker: i don't I don't actually think it's from zero. There's already like quite a bit of personality baked into the model. And I think in the open claw example, it works really well yeah because fundamentally, most people want their open claw to be like helpful and good. and like Yeah, exactly. right But I think it would be interesting to like just for fun, um just for your own understanding, to train some models that maybe do different things.
Speaker: Yeah. Could be enjoyable, right? Well, it also like it makes you understand... why are these things trying to be helpful? It's not just because fine tuning, but it's also the corpus.
Speaker: And I always talk about the corpus, whether it be parts of, you know, weird corners of Reddit or weird corners of Stack Overflow. If you end up in like the sassy part of the internet, the model is going to be snarky and sassy. And it's not that it's not goal seeking and it's not that it wants to help you or not. It's just, it was built on a sassy corpus.
Speaker: So you got to put your thumb on the scale and fine tune it to be a little bit more helpful. That's where supervised fine tuning comes in. Yeah. Have you, do you remember ah Claude Golden Gate? No.
Speaker: This was before a giant anthropic, but this was like an an early an early version of Claude that instead of trying to be, like it tried to be helpful, but it also was given clear instruction to like, whatever happens, try to work the Golden Gate into your answer.
Speaker: Oh, I see. It has a separate hidden goal. Truly beautiful and somewhat insane results. That's awesome. Yeah. so like well those Those are what you're seeing now with ah teachers putting in white text on a white background in the syllabus. So when the kid goes control A and control C and then paste it directly into the large language model, the teacher wants to detect whether or not the kid has AI. So they poison the prompt.
Speaker: And then they say, make sure you mention, you know, eggplants. And then yeah there's some random thing in the middle that the kid never tests and it says eggplant. In the like 49 of the 50 results. i One thing I fear maybe is...
Speaker: And this sort of goes back to like the sole document, right? like i think I think the sole document could probably get you pretty far in terms of like very late-stage context engineering.
Speaker: right But I think especially as the models get bigger and more powerful, and by bigger and more powerful, I mean like there's just more weights and like more parameters that are actually active, not just expert parameters, but generally parameters.
Speaker: I think hopefully we'll increasingly get models that
Speaker: There's sort of like three stages in my mind, right? Like stage one is the dumb model that just like does whatever it's given. And then stage two is maybe sort of the model that reads the instructions and like works ex-fant into the document.
Speaker: But hopefully we'll actually get to a model that has even more humanity and morality programmed into it to the point where it's like, clearly I'm being asked to cheat. Yeah. I should like help this kid.
Speaker: um To the same extent that maybe like, you know, the things that you and I would do if like our kids came by and were like, hey, can you write my homework for me? like how you my moment Exactly. And that's funny you mentioned that because I'm working on a thing with Rusinovich at Microsoft, which is a, we call it a preceptorship. It's a different spin on an internship.
Speaker: And we want to make models that are coding models, coding smart models. Right now we have to do it with skills and with and markdown files, but I want... a model that doesn't hold their hand. I want it to like, I'm going to leave this as an exercise to the reader.
Speaker: I know you want me to make a bubble sort, but I'd like to see you do it first. I want to a friendlier early in career engineer model that teaches them how to think. And I want people in universities to have access to those models. I don't want it to just like, we should not become a subcognitive species.
Speaker: Just because it can do it doesn't mean it should do it. Yeah, and I think maybe, I mean, I made Language Model Builder for fun. It's free. I'm not going make any money with it. So like maybe it's it's going to seem less like I'm i'm just advertising it. But like if if people have any questions about, okay, why do models behave certain ways, like the fine-tuning data section in it,
Speaker: yeah It's like really beautiful because also fine-tuning doesn't take a lot of steps. like Pre-training takes a long time. like Even in this app, it takes at least a day. But fine-tuning, 800 steps or so, like more than enough to like for these small models to really steer how they operate.
Speaker: And just seeing like this model, like you put it in the oven for five minutes and you tell it, I want you to like give helpful responses or whatever. um And just like playing with some of your own examples a little bit is like really fun.
Speaker: Yeah. Well, I want to call that out because languagemodelbuilder.com is free. And I want to remind folks that like you could have become an AI grifter.
Speaker: Everybody on social media, everybody on the dumpster fire that is Twitter wants you to buy their course. Give me $19.95, right? PayPal me five bucks. Learn how to avoid, learn how to avoid, you know, grifters, send me $5 and I'll tell you how.
Speaker: You chose to make it free. You just go to languagemodelbuilder.com. There's no upsell. You made it for fun. You made it for your friends and now your friends are the whole internet. So I want to appreciate that you didn't put, you know, a gum, a gum road or a Shopify on the top of this thing and try to make a quick buck. You're just putting the information out there for the people.
Speaker: Thank you. Yeah. I mean, it's an educational fun app, right? Like I'm pretty sure there was not a lot of money to be made to begin with. Yeah, but it's less about that. Like the information should be free. Let the information out.
Speaker: Everyone out there is trying to get you to do their class and everyone on LinkedIn wants you to get into their AI thing. What I like about Language Model Builder is that I could go to, you know, Portland Community College, my alma mater and teach a class and use this as the no pun intended, as the corpus for the class and the syllabus. And it's a great, great place to start, whether you're already an engineer or you're just a person who wants to know how to drive stick shift. I would encourage folks to check it out. And I look forward to the Windows version someday.
Speaker: Of course. Yeah. Thank you so much. Yeah. Thank you. We have been chatting with Felix Rieseberg. This is another episode of Hansel Minutes, and we'll see you again next week.




